system
The system addresses the inefficiencies in conventional biological observation by automating data processing from multiple cameras, enhancing species and individual identification, and improving accuracy through continuous learning and user-assisted data input.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2026-03-16
AI Technical Summary
Conventional biological observation systems face challenges in accurately measuring the number of species and inhabitants in an area, particularly in analyzing large amounts of video data from multiple observation points and performing detailed species and individual identification, with manual data processing being inefficient and limiting accuracy improvement.
A system that includes video acquisition from multiple cameras, automated organism detection and identification, data storage, registration of unidentified organisms for model retraining, and biodensity calculation, with user interface for inputting additional information to enhance accuracy over time.
Enables efficient and accurate species and individual identification, continuous learning, and improved biodensity calculations, allowing for precise measurement and analysis of organism populations.
Smart Images

Figure 2026047970000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In conventional biological observation systems, it has been difficult to accurately measure the number of species and the number of inhabitants of organisms in an area. In particular, it has been a problem to simultaneously analyze a huge amount of video data obtained from multiple observation points and perform detailed species identification including individual identification. In addition, in order to improve the accuracy of video analysis, it is necessary to continuously learn data of organisms that could not be identified and improve the performance of the system. However, in conventional systems, this process is often performed manually, which is inefficient.
Means for Solving the Problems
[0005] The present invention solves the above problems by providing a system that includes means for acquiring video footage from multiple cameras installed at specific observation points, means for analyzing the acquired video footage to detect and identify organisms in the footage, means for storing data of detected and identified organisms in a database, means for registering data of unidentified organisms as training data and retraining an AI model, and means for calculating the organism density for each observation point and generating a report. Furthermore, the present invention aims to improve identification accuracy by providing the generated report to the user through a user interface and allowing the user to input information on unidentified organisms. In addition, the present invention includes means for uniquely identifying each individual organism, in addition to species identification of organisms detected in the footage, thereby enabling detailed observation of each individual.
[0006] An "observation point" refers to a specific location where observation equipment, such as cameras, is installed to monitor the habitat conditions of living organisms.
[0007] A "camera" is a device used to acquire video data and is a device that plays a role in recording the movements of living organisms.
[0008] "Video footage" refers to a series of still images captured by a camera, which record the dynamics of living organisms over time.
[0009] "Analysis" refers to the data processing process performed to extract useful information from acquired video footage data.
[0010] "Living organisms" refers to all species of organisms that are observed, such as animals and insects.
[0011] "Detection" refers to the process of finding a specific organism within video data.
[0012] "Identification" is the process of identifying the type of organism detected and classifying it based on its unique characteristics.
[0013] A "database" refers to an information management system used to centrally manage and store analysis results and training data.
[0014] "Training data" refers to training data used to improve the accuracy of AI models, and includes data that contains information about organisms that are difficult to identify.
[0015] An "AI model" refers to an algorithm or data model that uses artificial intelligence to analyze video data and detect and identify living organisms.
[0016] "Retraining" refers to the process of training an AI model again using new training data in order to improve the model's accuracy.
[0017] "Biodensity" is an indicator that expresses the number of organisms at a specific observation point in units per area.
[0018] A "report" refers to a document that summarizes observation results and analytical data, including biodensity and details of detected organisms.
[0019] A "user interface" refers to the interface through which a user accesses a system and displays and manipulates the results.
[0020] "Individual identification" refers to the process of uniquely identifying each individual organism within the same species. [Brief explanation of the drawing]
[0021] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4]This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0022] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0023] First, let's explain the terminology used in the following explanation.
[0024] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).
[0025] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0026] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0027] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0029] [First Embodiment]
[0030] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0031] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0032] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0033] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0034] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0036] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0037] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0038] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0039] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0040] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0041] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0042] This invention relates to a system that acquires video footage from multiple cameras installed at a specific observation point, analyzes the acquired footage to identify biological species and individuals, and calculates biological density. The system has a configuration centered around a server, terminals, and users, and operates as follows.
[0043] Data collection
[0044] The server periodically acquires video footage from multiple cameras installed at observation points and stores it in a central database. For example, it acquires video every hour from observation points A and B in a mountainous area.
[0045] Video analysis
[0046] The server uses an AI model to analyze the acquired video footage. This model includes object detection algorithms using deep learning (e.g., YOLO and Mask R-CNN). The server uses this to detect living organisms in the video, identify their species, and identify the characteristics of each individual. For example, it can identify three deer and two rabbits from the video of observation point A.
[0047] Database update
[0048] The server stores data on detected and identified organisms in a central database. For organisms that are not identified, the video data is stored and used later as training data. Based on the data of newly identified organisms, the AI model is retrained to improve the accuracy of the analysis. For example, a newly discovered insect could be registered in the database.
[0049] density calculation
[0050] The server calculates the biodensity for each observation area based on the analysis results. The calculation results are compiled into a report and stored in a central database. For example, at observation point A, the deer density is 1.5 individuals / m³. 2 Rabbit density is 1 rabbit / m 2 Record this result in the report.
[0051] User Interface
[0052] The terminal provides an interface for users to access the system and view analysis results and biodensity reports. Users can view results for specific observation points and input information about unidentified organisms. This input information is used to train the AI model for the next time. For example, a user might view results for observation point A and supplement the information with unidentified organisms to facilitate learning.
[0053] As described above, this system can efficiently perform a series of processes including species and individual identification within the observation area, density calculation, and user-assisted information supplementation. This makes it possible to accurately measure the number of species and populations of organisms and to improve the accuracy of the analysis year after year.
[0054] The following describes the processing flow.
[0055] Step 1: Data Collection
[0056] The server periodically collects video footage from multiple cameras installed at the observation site.
[0057] The server stores the collected video data in a temporary storage folder and then saves it as a backup in the central database.
[0058] For example, video footage is collected from observation points A and B every hour and stored in a central database.
[0059] Step 2: Video Analysis
[0060] The server reads unanalyzed video data from the central database.
[0061] The server loads a pre-trained deep learning model (e.g., YOLO or Mask R-CNN) and analyzes the video data.
[0062] The server detects organisms in the video and assigns a bounding box, species label, and confidence score to each organism. For example, it identifies three deer and two rabbits from video footage of observation point A.
[0063] Step 3: Individual Identification
[0064] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species.
[0065] The server assigns a unique ID to each individual and identifies them by considering their distinctive physical characteristics (patterns, body size, etc.). For example, if one of three deer has a different pattern, the server will use that as the basis for individual identification.
[0066] Step 4: Database Update
[0067] The server stores the species and individual identification information of the identified organisms in a database.
[0068] The server adds data of unidentified organisms to an unidentified list and saves the video footage for later use as training data. For example, it adds newly discovered insects to the unidentified list.
[0069] Step 5: Register training data and retrain the AI model
[0070] The user checks the unidentified list and enters the species name and information of the newly identified organism.
[0071] The server updates the training data based on new biological information provided by the user and prepares to retrain the AI model.
[0072] The server retrains to improve the recognition accuracy of the AI model.
[0073] Step 6: Density Calculation
[0074] The server compiles the number of organisms in each observation area based on the analysis results.
[0075] The server uses the aggregated number of organisms to determine the biodensity (number of individuals / m²) for each area. 2 Calculate the density of deer at observation point A. For example, the density of deer at observation point A is 1.5 deer / m². 2 Rabbit density is 1 rabbit / m 2 This is the result.
[0076] Step 7: Report Generation
[0077] The server generates reports for each observation area based on the calculated biodensity.
[0078] The server saves the generated reports to a central database.
[0079] Step 8: Displaying the User Interface
[0080] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user.
[0081] Users review the observation results through the interface and input information about unidentified organisms as needed. This information is then used to train the AI model for the next time.
[0082] Through the steps described above, this system can efficiently measure the number of species and population size of organisms, and its accuracy can be improved year by year.
[0083] (Example 1)
[0084] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0085] In nature observation and environmental monitoring, efficiently analyzing large amounts of video data acquired from cameras installed at specific observation points and accurately identifying species and individuals of organisms is crucial. However, conventional methods have mainly relied on manual analysis, which is labor-intensive, time-consuming, and suffers from accuracy issues. Furthermore, there has been no mechanism to effectively utilize data on unidentified organisms and continuously improve the accuracy of the analysis. This has resulted in limitations in the accuracy and reusability of observation data, making accurate calculation of biodensities difficult.
[0086] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0087] In this invention, the server includes means for acquiring video data from multiple image acquisition devices installed at observation sites, means for storing the acquired video data in a data storage device, means for inputting the stored video data into an analysis device for analysis and automatically detecting and identifying organisms in the video, means for storing the data of identified organisms in a database, means for registering the data of unidentified organisms as training data and retraining the recognition model device, and means for calculating the organism density for each observation site and generating report data. This enables efficient and highly accurate analysis of large amounts of video data, and by continuously improving the analysis accuracy, accurate calculation of organism density becomes possible.
[0088] An "image acquisition device" is a device installed at a specific observation point to acquire video data of the surrounding area.
[0089] "Video data" refers to video and still image data acquired by an image acquisition device.
[0090] A "data storage device" is a device used to store acquired video data. Examples include databases and storage servers.
[0091] An "analysis device" is a device that analyzes acquired and stored video data to automatically detect and identify biological species and individuals. It primarily uses AI models and deep learning algorithms.
[0092] "Living organisms" refers to plants, animals, and other organisms detected at the observation site.
[0093] A "database" is a system for managing and storing data on identified organisms.
[0094] "Data on organisms that could not be identified" refers to video data of organisms that the analysis device could not automatically identify.
[0095] "Training data" refers to data used to retrain an AI model.
[0096] A "recognition model device" is a device that operates and manages AI models used for identification.
[0097] "Biodensity" refers to the number of organisms per unit area at a specific observation point.
[0098] "Report data" refers to data compiled in the form of a report, which includes the results of biodensity calculations and analyses.
[0099] A "user interface" is the interface through which a user accesses a system and views analysis results and reports.
[0100] "Users" refer to individuals who access the system, view analysis results, or input information about unidentified organisms.
[0101] This invention relates to a system that acquires video data from multiple image acquisition devices installed at an observation site, analyzes the acquired video data to identify biological species and individuals, and further calculates biological density. The system has a configuration centered around a server, terminals, and users, and operates as follows.
[0102] Data collection
[0103] The server periodically acquires video data from multiple image acquisition devices (e.g., cameras) installed at observation sites and stores it in a central database. High-resolution cameras are suitable for use. For example, a configuration could be set up to acquire video data from observation sites A and B every hour. As a specific example, video taken at observation site A at 10:00 AM would be saved as "A_20231001_10.mp4".
[0104] Video analysis
[0105] The server analyzes the acquired video data using an analysis device (e.g., a high-performance GPU server). This analysis device is equipped with an object detection algorithm using deep learning (e.g., YOLO, Mask R-CNN). The server uses this AI model to detect living organisms in the video, identify each species, and analyze the characteristics of each individual. For example, one might identify three deer and two rabbits from video footage of observation point A. In this case, software such as Python and TENSORFLOW® is used.
[0106] Database update
[0107] The server stores data on detected and identified organisms in a database. For organisms that are not identified, their video data is saved as training data and later used to retrain the AI model. The AI model is retrained based on data of newly identified organisms to improve the accuracy of the analysis. For example, a newly discovered insect might be registered in the database, and this data could be used to improve the accuracy of the model.
[0108] density calculation
[0109] The server calculates the biodensity at each observation point based on the analysis results. These results are compiled into a report and stored in a database. For example, at observation point A, the deer density is 1.5 individuals / m³. 2 Rabbit density is 1 rabbit / m 2 This is recorded in the report. R or the Pandas library in Python are used for the calculations.
[0110] User Interface
[0111] The terminal provides an interface for users to access the system and view analysis results. This interface runs on a web browser and has the functionality to view results for a specific observation location. Users can input information about unidentified organisms, and this input information is used for the next AI model retraining. For example, a user might view the results for observation location A and supplement the information about unidentified organisms to accelerate learning. The interface is implemented using React, with Node.js and Express used for the backend.
[0112] As a concrete example, the following prompt can be input to the generating AI model:
[0113] Please describe in natural language the process of a system that acquires video footage from multiple cameras installed at observation sites, uses an AI model to detect organisms in the footage, and identifies the characteristics of species and individuals. Please explain in detail what kind of data processing and calculations are performed using specific hardware (e.g., high-resolution cameras, GPU servers) and software (e.g., Python, TensorFlow, YOLO, React, Node.js). For example, please explain how and from where data is acquired, how it is analyzed, how the analysis results are stored, how users can view the results, and how they can input information about unidentified organisms.
[0114] This invention enables efficient execution of a series of processes, including species and individual identification within an observation area, density calculation, and user-assisted information supplementation. This allows for accurate measurement of the number of species and populations, and improves the accuracy of the analysis year after year.
[0115] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0116] Step 1:
[0117] Data collection
[0118] The server acquires video data from multiple image acquisition devices installed at observation sites. Specifically, it acquires video every hour from cameras installed at observation sites A and B. The input is video data from the image acquisition devices, and the output is the acquired video data. For example, video taken at observation site A at 10:00 AM is saved as "A_20231001_10.mp4". This data is transferred to a central database via the high-speed internet.
[0119] Step 2:
[0120] Saving video data
[0121] The server stores the acquired video data in a data storage device. The server uses a MySQL® database to save the video data in an appropriate format. The input is the video data acquired in step 1, and the output is the video data stored in the database. For example, "A_20231001_10.mp4" is stored in the MySQL database.
[0122] Step 3:
[0123] Video analysis
[0124] The server inputs stored video data into an analysis device (high-performance GPU server) for analysis. The server utilizes deep learning-based object detection algorithms (e.g., YOLO, Mask R-CNN). The input is video data read from a database, and the output is detected and identified biological data. For example, three deer and two rabbits are identified from video footage of observation point A. This process uses Python and TensorFlow.
[0125] Step 4:
[0126] Database update
[0127] The server stores the analyzed biological data in a database. Specifically, it stores species and individual data of detected and identified organisms. The input is the analysis result data, and the output is the biological data stored in the database. Video data of organisms that were not identified is also stored and used later as training data. For example, data for a newly discovered organism is stored as "species: unidentified, quantity: 1".
[0128] Step 5:
[0129] Model Retraining
[0130] The server retrains the AI model using data on organisms that were not identified. The server uses TensorFlow to retrain the AI model. The input is the new training data, and the output is the retrained AI model. After retraining, an evaluation test is run to confirm the improvement in accuracy.
[0131] Step 6:
[0132] density calculation
[0133] The server calculates the biological density for each observation point based on the analysis results. The server performs the calculation using the R or Python Pandas library. The input is the analyzed biological data, and the output is the calculation result of the biological density. For example, in the report of observation point A, "Density of deer = 1.5 per m 2 , Density of rabbit = 1 per m 2 " is obtained.
[0134] Step 7:
[0135] Report generation and storage
[0136] The server generates a report based on the calculation results and stores it in the database. The generated report is saved in PDF format. The input is the result data of the density calculation, and the output is the generated report. For example, it is saved as "report_A_20231001.pdf".
[0137] Step 8:
[0138] Providing a user interface
[0139] The terminal provides an interface for users to access the system and view the analysis results and reports. The interface operates on a web browser and is implemented using React. The input is the access request from the user, and the output is the display of the analysis results and the function to download the report. Users can view the results of a specific observation point and input information about unidentified organisms.
[0140] Step 9:
[0141] Reflecting user input
[0142] Users input information about unidentified organisms through the user interface. This information is stored in the database and used for the next retraining of the AI model. The input is the information about unidentified organisms input by the user, and the output is the updated training data.
[0143] summary
[0144] Thus, this system efficiently executes a series of processes, from data collection and analysis to database updates, density calculations, and user interface input. By using specific hardware and software, accurate identification of species and individuals, as well as density calculations, are possible, and continuous learning and accuracy improvement are achieved.
[0145] (Application Example 1)
[0146] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0147] Conventional animal observation systems focus on identifying species and individuals, but lack the ability to identify specific individuals in real time and monitor their density and intrusion status. This has resulted in inefficient security management within facilities, leading to delays in detecting suspicious individuals and issuing alarms. Furthermore, the lack of a system that simultaneously performs both animal observation and person identification has made integrated data management difficult. This invention aims to improve facility security levels by enabling real-time identification of individuals alongside animal observation.
[0148] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0149] In this invention, the server includes means for acquiring video footage from multiple cameras, means for analyzing the acquired video footage to detect and identify animals, means for storing data of detected and identified animals in a storage device, means for registering data of unidentified animals as training data and retraining an AI model, means for calculating animal density for each observation point and generating a report, means for analyzing the acquired video footage to identify specific individuals and their density in real time, means for storing data of detected and identified individuals in a database, means for calculating the density of individuals in each area based on the analysis results, means for inputting information on unidentified animals, and means for issuing alarms based on identified individuals. This makes it possible to efficiently perform both animal observation and security monitoring.
[0150] A "filming device" is a device installed at a specific observation point to capture video footage.
[0151] "Video footage" refers to a series of image data acquired from multiple recording devices.
[0152] "Identification" is the process of distinguishing individual animals or people as specific species from acquired video footage.
[0153] "Detection" is the process of finding objects in acquired video footage.
[0154] A "storage device" is hardware or software used to store data.
[0155] A "database" is a system for efficiently storing and managing structured data.
[0156] An "AI model" is an artificial intelligence program created based on machine learning algorithms and trained to perform a specific task.
[0157] "Real-time" means processing and analyzing acquired data immediately and providing the results instantly.
[0158] "Density" is an indicator that shows the number of animals or people present within a given observation area.
[0159] A "report" is a document that summarizes observational data and analysis results.
[0160] An "alarm" is a warning signal that is issued when specific conditions are met.
[0161] A "domain" refers to a specific area or place that is the subject of observation or surveillance.
[0162] A "user interface" is software that provides a means for a system and a user to interact.
[0163] "Training data" refers to the dataset used to train an AI model.
[0164] "Retraining" is the process of retraining an existing AI model using new data.
[0165] This invention relates to a system that acquires video footage from multiple cameras installed at observation points, analyzes the acquired footage to detect and identify animals and people, and uses the results for security management. The system has a configuration centered around a server, terminals, and users, and operates as follows.
[0166] Data collection
[0167] The server periodically acquires video footage from multiple cameras installed within the facility and stores it in a central database. For example, it might acquire video footage every hour from a camera installed in a specific research facility.
[0168] Video analysis
[0169] The server uses deep learning-based object detection algorithms (e.g., YOLO or Mask R-CNN) to analyze the acquired video footage. This algorithm is used to identify animals and specific individuals within the video. For example, it can identify three deer, two rabbits, and a specific person from footage from a certain camera.
[0170] Database update
[0171] The server stores data on detected and identified animals and people in a central database. For animals that are not identified, the video data is stored as training data and later used to retrain the AI model.
[0172] density calculation
[0173] The server calculates the density of animals and people in each observation area based on the analysis results. The calculation results are compiled into a report and stored in a central database. For example, at a specific observation point, the deer density is 1.5 individuals / m². 2 Rabbit density is 1 rabbit / m 2 The density of people is 0.2 people / m². 2 Record this result in the report.
[0174] Alarm and monitoring
[0175] The server has a means of issuing real-time alarms based on identified individuals. For example, if an unauthorized person enters a specific authorized area, an alarm is automatically issued and the responsible person is notified.
[0176] User Interface
[0177] The terminal provides an interface for users to access the system and view analysis results and animal and human density reports. Users can view results for specific observation locations and input information on unidentified animals. This input information will be used to train the next AI model.
[0178] Hardware and software to be used
[0179] Hardware:
[0180] Recording equipment: Surveillance cameras installed within the facility (e.g., typical network cameras)
[0181] Server: A computer that runs the database and deep learning models (e.g., typical server hardware).
[0182] software:
[0183] OpenCV: Used to capture and save camera footage.
[0184] PyTorch: A deep learning framework
[0185] SQLite: Database Administration
[0186] Flask: Building User Interfaces
[0187] Examples of processes and prompt statements
[0188] Specific example:
[0189] Cameras installed in a research facility capture video in real time to detect intrusions by specific individuals. The detection results are stored in a central database and notified to security personnel in real time.
[0190] Example of a prompt:
[0191] Analyze the surveillance camera footage within the facility to identify specific individuals and their density in real time.
[0192] For example, it checks whether a man wearing a white coat is in an authorized area.
[0193] As described above, this system can improve facility security by efficiently performing a series of processes including animal species and individual identification within the observation area, density calculation, real-time person identification and density calculation, alarm output, and user-assisted information supplementation.
[0194] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0195] Step 1:
[0196] The server acquires video footage from multiple cameras installed within the facility. The cameras periodically capture video data and send it to the server. The input is video data from the cameras, and the output is a saved video file.
[0197] Step 2:
[0198] The server analyzes the acquired video footage using deep learning-based object detection algorithms (e.g., YOLO or Mask R-CNN). Here, animals and people are detected within the video. The input is the video file acquired in step 1, and the output is a list of detected animals and people along with their location information.
[0199] Step 3:
[0200] The server stores data on detected and identified animals and people in a central database. Attribute information is also stored for each identified individual. The input is the list and location information obtained in step 2, and the output is the data stored in the database.
[0201] Step 4:
[0202] The server registers data of animals that were not identified as training data and uses it to retrain the AI model. The input is video data of animals that were not identified in step 2, and the output is the updated AI model.
[0203] Step 5:
[0204] The server calculates the density of animals and people in each observation area. Here, it calculates the number of animals and people present in a specific area and determines their density. The input is the data saved in step 3, and the output is the calculated density information.
[0205] Step 6:
[0206] The server generates a report for each observation area based on the calculation results. The report includes animal species and density, human density, and all detected data. The input is the density information obtained in step 5, and the output is the generated report.
[0207] Step 7:
[0208] The terminal provides the generated report to the user through a user interface. The user can review the results for a specific observation point and supplement information on unidentified animals. The input is the report generated in step 6, and the output is the analysis results displayed to the user.
[0209] Step 8:
[0210] The server identifies specific individuals and issues alarms in real time. For example, if an unauthorized person enters a specific authorized area, an alarm is automatically triggered. The input is data on the person detected in step 2, and the output is a record of the alarms that were triggered.
[0211] This series of processing steps enables the creation of a system that efficiently performs both animal observation and security monitoring within the facility.
[0212] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0213] This invention is a system that acquires video footage from multiple cameras installed at a specific observation point, analyzes the data to identify species and individuals, and calculates biodensity. Furthermore, it features an emotion engine that recognizes user emotions to improve the user experience. This system has a configuration centered around a server, terminals, and users, and operates as follows.
[0214] Data collection
[0215] The server periodically collects video footage from multiple cameras installed at observation points and stores it in a central database. For example, it collects video from observation points A and B every hour and stores it in the central database.
[0216] Video analysis
[0217] The server uses an AI model to analyze the acquired video footage. This model includes object detection algorithms using deep learning (e.g., YOLO and Mask R-CNN). The server uses this to detect living organisms in the video, identify their species, and identify the characteristics of each individual. For example, it can identify three deer and two rabbits from the video of observation point A.
[0218] Individual identification
[0219] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species. It assigns a unique ID to each individual and identifies them based on distinctive physical characteristics (patterns, body size, etc.). For example, if one of three deer has a different pattern, it will be identified based on that.
[0220] Database update
[0221] The server stores the species and individual identification information of identified organisms in a database. It adds data of unidentified organisms to an unidentified list and saves their images for later use as training data. For example, it adds newly discovered insects to the unidentified list.
[0222] Registration of training data and retraining of the AI model
[0223] The user reviews the unidentified list and enters the species name and information of newly identified organisms. This information is used to train the AI model for the next time. The server updates the training data based on the new biological information provided by the user and retrains the AI model. This improves the accuracy of the next analysis.
[0224] density calculation
[0225] The server aggregates the number of organisms in each observation area based on the analysis results, and calculates the biodensity (number of individuals / m³) for each area. 2 Calculate the density of deer at observation point A. For example, the density of deer at observation point A is 1.5 deer / m². 2 Rabbit density is 1 rabbit / m 2 This is the result.
[0226] Report generation
[0227] The server generates reports for each observation area based on the calculated biodensity. These generated reports are then stored in a central database.
[0228] User Interface Display
[0229] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user. Users can review results for specific observation points and input information about organisms that were not identified. This input information is used to train the AI model for the next analysis, enabling more accurate analysis.
[0230] Combination of emotional engines
[0231] The server uses an emotion engine to recognize the user's emotions. Based on this emotion data, it can customize how analysis results and reports are displayed. For example, if a user is stressed, the frequency and level of detail of information provided can be adjusted to reduce their burden. Furthermore, the emotion engine can also analyze the context and emotions of unidentified organisms entered by the user, enabling automatic completion and correction.
[0232] Specific example
[0233] For example, based on video footage acquired from observation point A, the server identifies three deer and two rabbits and assigns a unique ID to each individual. Later, when the user uses a device to view the analysis results, the emotion engine recognizes the user's interests and stress levels and customizes the displayed content. When the user discovers a new, unidentified insect and enters its information, the emotion engine analyzes the context and emotions of the input and automatically completes the information.
[0234] This system enables efficient measurement of the number of species and population size of organisms, improving analysis accuracy and optimizing the user experience.
[0235] The following describes the processing flow.
[0236] Step 1: Data Collection
[0237] The server periodically collects video footage from multiple cameras installed at the observation site.
[0238] The server stores the collected video data in a temporary storage folder and simultaneously saves it as a backup in the central database.
[0239] For example, video footage is collected from observation points A and B every hour and stored in a central database.
[0240] Step 2: Video Analysis
[0241] The server reads unanalyzed video data from the central database.
[0242] The server loads a pre-trained deep learning model (e.g., YOLO or Mask R-CNN) and analyzes the video data.
[0243] The server assigns bounding boxes, species labels, and confidence scores to organisms detected in the video. For example, it identifies three deer and two rabbits from video footage of observation point A.
[0244] Step 3: Individual Identification
[0245] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species.
[0246] The server assigns a unique ID to each individual and identifies them by considering their distinctive physical characteristics (patterns, body size, etc.). For example, if one of three deer has a different pattern, the server will use that as the basis for individual identification.
[0247] Step 4: Database Update
[0248] The server stores the species and individual identification information of the identified organisms in a database.
[0249] The server adds data of unidentified organisms to an unidentified list and saves the video footage for later use as training data. For example, it adds newly discovered insects to the unidentified list.
[0250] Step 5: Register training data and retrain the AI model
[0251] The user checks the unidentified list and enters the species name and information of the newly identified organism.
[0252] The server updates the training data based on the new biological information provided by the user and prepares to retrain the AI model.
[0253] The server performs retraining to improve the identification accuracy of the AI model.
[0254] Step 6: Density calculation
[0255] The server aggregates the number of organisms for each observation area based on the analysis results.
[0256] The server calculates the organism density (number of individuals / m 2 ) for each area based on the aggregated number of organisms. For example, at observation point A, the deer density is 1.5 individuals / m 2 , and the rabbit density is 1 individual / m 2 .
[0257] Step 7: Report generation
[0258] The server generates a report for each observation area based on the calculation results of the organism density.
[0259] [[ID=3The server uses an emotion engine to recognize the user's emotions. For example, it analyzes sensor data from webcams and microphones to evaluate the user's facial expressions and tone of voice.
[0265] The server customizes how analysis results and reports are displayed based on emotional data. For example, if a user is experiencing stress, it adjusts the frequency and level of detail of information provided to reduce the user's burden.
[0266] Users also use the sentiment engine to analyze the context and emotions of the unidentified organisms they input. This allows for simple corrections and automatic completion, ensuring that accurate information is reflected in the database.
[0267] Through the steps described above, this system can efficiently measure the number of species and population size of organisms, improve analysis accuracy year after year, and optimize the user experience.
[0268] (Example 2)
[0269] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0270] Conventional biological observation systems can identify species when analyzing video footage from observation sites, but they lack the means to uniquely identify individual organisms. Furthermore, they fail to provide information that takes into account the user's emotional state, resulting in an unoptimized user experience. Therefore, improvements in both observation accuracy and user experience are needed.
[0271] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0272] In this invention, the server includes means for acquiring video data from a plurality of imaging devices installed at an observation point, means for analyzing the acquired video data to detect and identify organisms in the video, means for storing the data of the detected and identified organisms in a storage device, means for registering the data of organisms that have not been identified as learning data and retraining an artificial intelligence model, means for calculating the organism density for each observation point and generating a report, means for uniquely identifying each individual organism using the artificial intelligence model, and means for recognizing the user's emotion and customizing the analysis results and report based on the data. Thereby, the observation accuracy of organisms is improved, and it becomes possible to optimize the user experience.
[0273] An "imaging device" is a device installed at an observation point for acquiring video data.
[0274] "Video data" refers to video information acquired from a plurality of imaging devices.
[0275] A "server" is a computing device for analyzing video data and detecting and identifying organisms.
[0276] A "storage device" is a data storage for storing the data of detected and identified organisms.
[0277] An "artificial intelligence model" is a general term for machine learning algorithms used for detecting and identifying organisms and uniquely identifying individuals.
[0278] "Retraining" is a learning process for improving the accuracy of an artificial intelligence model using the data of organisms that have not been identified.
[0279] "Organism density" refers to the number of organisms per unit area at a specific observation point.
[0280] A "report" is a document summarizing analysis results, organism density, etc.
[0281] "User emotions" refer to the psychological state a user experiences while using a system, such as interest or stress.
[0282] "Customization" refers to adjusting how analysis results and reports are displayed based on user sentiment data.
[0283] This invention is a system that acquires video data from multiple cameras installed at a specific observation point, analyzes that data to identify species and individuals, and calculates biodensity. Furthermore, it is characterized by its ability to improve the user experience by incorporating an emotion engine that recognizes the user's emotions. This system has a configuration centered on a server, terminals, and users, and operates as follows.
[0284] Data collection
[0285] The server periodically collects video data from multiple cameras installed at observation sites and stores it in a central database. For example, it collects video data from observation sites A and B every hour and stores it in the central database. It downloads and manages video streams using HTTP requests.
[0286] Video analysis
[0287] The server loads a pre-trained deep learning model (such as YOLO or Mask R-CNN) to analyze the acquired video data. This model is used to analyze each frame of the video, and an object detection algorithm is applied. The GPU is used to perform model inference at high speed, detecting the location and species of organisms in the video. For example, analyzing video from observation point A can detect three deer and two rabbits in each frame.
[0288] Individual identification
[0289] The server applies additional deep learning models and clustering algorithms to each detected organism to identify individuals within the same species. Each individual is assigned a unique ID and identified based on characteristics such as markings and body size. For example, if one of three deer has a different marking, it is identified as "Deer 1," "Deer 2," and "Deer 3" based on that.
[0290] Database update
[0291] The server updates the database with species and individual identification information for identified organisms. It uses SQL queries to insert the identification information into the appropriate tables. Data for unidentified organisms is added to an unidentified list, and their videos are saved separately. For example, a newly discovered insect is added to the unidentified list, and its video is saved to cloud storage.
[0292] Registration of training data and retraining of the AI model
[0293] The user reviews the unidentified list and enters the species name and other information of newly identified organisms. This operation can be performed through a form in the web application. This information is sent to the server in JSON format. The server incorporates the new organism data provided by the user into the training database and generates a new training set. The AI model is retrained using a GPU cluster to improve accuracy.
[0294] density calculation
[0295] The server aggregates the number of organisms in each observation area based on the analysis results and calculates the biodensity. It obtains area information for each observation point from the database and calculates the density by dividing the number of individuals by the area. For example, using data obtained from the scan of observation point A, it calculates that "the density of deer is 1.5 individuals / m²". 2 "The rabbit density is 1 rabbit / m²" 2 The calculation is as follows:
[0296] Report generation
[0297] The server generates reports for each observation area based on the calculated biodensity. A Python script is used to automatically generate these reports, which are then stored in a central database. These reports include density information and individual identification information for each species.
[0298] User Interface Display
[0299] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user. This interface is designed using front-end technologies such as JavaScript® and React. Users can view results for specific observation points and input information about organisms that were not identified. This input information is sent to the server and used for training the next AI model.
[0300] Combination of emotional engines
[0301] The server uses an emotion engine to recognize the user's emotions. It analyzes user operation logs and input data and applies algorithms to predict the emotional state. Emotional data is reflected in report displays and alert generation. For example, if the user is stressed, the frequency of information provided may be reduced or the display mode may be switched to a simplified mode. Even when the user enters information about an unfamiliar insect, the emotion engine analyzes the context and emotions and performs automatic completion or correction.
[0302] Specific example
[0303] For example, based on video data acquired from observation point A, the server identifies three deer and two rabbits and assigns a unique ID to each individual. Subsequently, when the user uses a device to view the analysis results, the emotion engine recognizes the user's interests and stress levels and customizes the displayed content. When the user discovers a new unidentified insect and enters its information, the emotion engine analyzes the context and emotions of the input and automatically completes the information.
[0304] Example of a prompt
[0305] Analyze the following video footage and identify the species of organisms shown. Also, identify each individual organism and update the database with your findings.
[0306] For organisms on the unidentified list, please retrain the AI model based on the information entered by the user to improve the identification accuracy for the next time.
[0307] Analyze user sentiment and customize how the analysis results and reports are displayed accordingly.
[0308] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0309] Step 1:
[0310] The server periodically collects video data from multiple cameras installed at observation sites. Its input is the video stream from the observation sites, obtained via HTTP requests. Specifically, it uses a scheduling function to access the cameras every hour and download the video data. The output is the storage of the acquired video data in a central database.
[0311] Step 2:
[0312] The server loads a pre-trained deep learning model (e.g., YOLO or Mask R-CNN) to analyze the acquired video data. The input is the stored video data. Specifically, each frame of the video is divided, and an object detection algorithm is applied. The output is the location information and species identification result of organisms in each frame. Model inference is performed at high speed using a GPU, and analysis is performed efficiently.
[0313] Step 3:
[0314] The server applies additional deep learning models and clustering algorithms to the detected organisms to identify individuals within the same species. The input is the location information and species identification results of the organisms obtained in step 2. Specifically, a unique ID is assigned to each individual, and features such as patterns and body size are analyzed. The output is the uniquely identified individual information. This makes it possible to make specific identifications such as "Deer 1," "Deer 2," and "Deer 3."
[0315] Step 4:
[0316] The server updates the database with the species and individual identification information of the identified organisms. The input is the unique individual information obtained in step 3. Specifically, it inserts this information into the appropriate table in the database using an SQL query. The output is the updated database information. Data of organisms that were not identified is added to the unidentified list, and their video data is saved to cloud storage.
[0317] Step 5:
[0318] The user reviews the unidentified list and enters the species name and other information of newly identified organisms. This input is provided through a web application form. This information is sent to the server in JSON format. The server incorporates this new data into the training database and generates a new training set. The output is an updated training data set and a new training set to be used for the next training session.
[0319] Step 6:
[0320] The server aggregates the number of organisms in each observation area based on the analysis results and calculates the biodensity. The input is the database information updated in step 4. Specifically, it obtains area information for each observation point from the database and calculates the density by dividing the number of individuals by the area. The output is the biodensity information for each observation area. For example, from the data for observation point A, it might say "the deer density is 1.5 individuals / m²". 2 "The rabbit density is 1 rabbit / m²"2 The calculation is as follows:
[0321] Step 7:
[0322] The server generates reports for each observation area based on biodensity. The input is the biodensity information obtained in step 6. A Python script is used to automatically generate the reports and save them to the central database. The output is the generated report. This report includes density information and individual identification information for each species.
[0323] Step 8:
[0324] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user. Input consists of reports and analysis results obtained from the server. The web page is designed using front-end technologies such as JavaScript and React. Output consists of analysis results and reports displayed to the user. The user can review results for specific observation areas and add information about organisms that were not identified.
[0325] Step 9:
[0326] The server uses an emotion engine to recognize the user's emotions. Input is the user's operation logs and input data. An algorithm that predicts emotional states is used to detect stress, interest, etc. The output is the recognized emotion information. Based on this, the frequency of information provision and the display method are automatically customized. Even when the user inputs information about an unknown insect, the emotion engine analyzes the context and emotions and automatically completes or corrects the information.
[0327] (Application Example 2)
[0328] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0329] There is a need to accurately understand customer behavior and emotions in stores and use that information to improve product placement and service quality, but current technology makes it difficult to fully achieve this. Furthermore, there is a lack of efficient methods for calculating and identifying biological densities. Therefore, more advanced data analysis technologies are needed to improve the efficiency of store operations and enhance the customer experience.
[0330] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0331] In this invention, the server includes means for acquiring video footage from multiple cameras installed at observation points, means for analyzing the acquired video footage to detect and identify organisms in the footage, means for storing data of the detected and identified organisms in a database, means for registering data of organisms that were not identified as training data and retraining an AI model, means for calculating the organism density for each observation point and generating a report, means for filming and analyzing customer behavior from multiple cameras installed in the store, means for analyzing the video footage to identify customers and assign a unique ID, and means for identifying emotional data using an emotion engine that analyzes customer emotions. This makes it possible to analyze customer behavior and emotions in detail, optimize store operations based on this analysis, and improve the customer experience.
[0332] An "observation point" is a specific area or location where the system acquires video footage and uses it for analysis.
[0333] A "camera" is a device used to capture video footage and collect that data.
[0334] A "server" is a central processing unit used to analyze collected video footage, store data, and retrain AI models.
[0335] "Video footage" is a media format composed of a series of still images that records the situation at an observation point or inside a store in real time.
[0336] "Living organisms" refers to species and individuals of animals, plants, insects, and other organisms present at the observation site.
[0337] "Identification" refers to the act of identifying a specific species, individual, or customer from analyzed video footage.
[0338] A "database" is an information system for systematically storing collected and analyzed data.
[0339] An "AI model" is an algorithm that uses deep learning or machine learning to perform data analysis.
[0340] "Retraining" is the process of improving the accuracy of an AI model using new training data.
[0341] "Biological density" is an index calculated by determining the number of organisms in a specific observation location per unit area.
[0342] A "report" is a quantitative and qualitative report generated based on the results of an analysis.
[0343] A "customer" is a person who engages in activities within a store.
[0344] An "emotion engine" is an analytical device and software that estimates emotions from a customer's facial expressions and body movements.
[0345] A "unique ID" is identification information that assigns a unique identifier to each identified organism or customer.
[0346] This invention is a system that acquires video footage from multiple cameras installed at specific observation points, analyzes the data to identify biological species and individuals, and calculates biodensity. It also aims to improve the user experience by incorporating an emotion engine that recognizes user emotions. Furthermore, by applying this system to physical stores and analyzing customer behavior and emotions, it aims to optimize store operations and improve the quality of customer service.
[0347] 1. Data Collection
[0348] The server periodically collects video footage from multiple cameras installed at observation points and stores it in a central database. For example, it collects video from observation points A and B every hour and stores it in the central database. It also periodically acquires video data from cameras installed inside the store. This allows for continuous monitoring of customer behavior.
[0349] 2. Video Analysis
[0350] The server uses the following AI model to analyze the acquired video footage.
[0351] Object detection algorithms such as YOLO (You Only Look Once) or Mask R-CNN
[0352] These models are used to detect organisms in video footage and identify their species and individual characteristics. Additionally, in in-store video footage, customers are identified, and each customer is assigned a unique ID.
[0353] 3. Individual Identification
[0354] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species or each customer. For example, it can identify individuals based on unique physical characteristics or body movements.
[0355] 4. Database Update
[0356] The server stores the species and individual identification information of identified organisms in a database. Unidentified data is added to an unidentified list, and its video is saved for later use as training data. The same applies to customers, storing customer behavior data and sentiment data in the database.
[0357] 5. Registering training data and retraining the AI model
[0358] Users can review the unidentified list and input the species name and information of newly identified organisms. Based on this, the server retrains the AI model. Similarly, for customer data, the model is retrained based on customer behavior patterns and sentiment data.
[0359] 6. Emotion analysis
[0360] The server uses an emotion engine to recognize the user's emotions. Based on this emotion data, it customizes how analysis results and reports are displayed. For example, if a user is feeling stressed in a store, the frequency and level of detail of information provided will be adjusted. It is also possible to analyze customer emotion data and provide appropriate services.
[0361] 7. Report generation and user interface display
[0362] The server generates reports for each observation area and store based on calculations of biodensity and customer behavior. These reports are stored in a central database and accessible to users via terminals. Users can review the analysis results and reports and input information about unidentified organisms and customers.
[0363] Specific example
[0364] For example, based on video footage acquired from observation point A, the server identifies three deer and two rabbits and assigns a unique ID to each individual. When the user then reviews the analysis results, the emotion engine recognizes the user's interests and stress levels and customizes the displayed content. Furthermore, when the user discovers a new unidentified insect and inputs its information, the emotion engine analyzes the context and emotions of the input and automatically completes the information.
[0365] Example of a prompt
[0366] "Please analyze the behavior of customers who spend extended periods of time in specific areas of the store (e.g., the cosmetics section), and based on the results of this analysis, including emotional data, please propose improvements."
[0367] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0368] Step 1:
[0369] Data collection
[0370] The server periodically collects video footage from multiple cameras installed at observation points and within the store. For example, the server collects video from observation points A and B every hour and stores it in a central database. It also acquires video data from cameras inside the store and stores it in the central database as well.
[0371] Input: Camera footage
[0372] Output: Video data stored in the central database
[0373] Step 2:
[0374] Video analysis
[0375] The server uses an AI model to analyze the acquired video footage. Specifically, it employs object detection algorithms such as YOLO and Mask R-CNN. This allows it to detect living beings and customers in the video and assign a unique ID to each individual or customer.
[0376] Input: Video data
[0377] Output: Identified individual / customer data (with unique ID)
[0378] Step 3:
[0379] Individual identification
[0380] The server further analyzes the detailed characteristics of identified organisms and customers using deep learning models and clustering algorithms. It uniquely identifies each individual or customer based on their physical characteristics and behavioral patterns.
[0381] Input: Identified individual / customer data
[0382] Output: Detailed data on individual customers (feature-based unique IDs)
[0383] Step 4:
[0384] Database update
[0385] The server stores data on identified organisms and customers in a database. Data that is not identified is added to an unidentified list, and this video is saved and used as training data later.
[0386] Input: Individual / customer detailed data, unidentified data
[0387] Output: Updated database
[0388] Step 5:
[0389] Registration of training data and retraining of the AI model
[0390] The user reviews the unidentified list and inputs information about newly identified organisms or customers. Based on this information, the server retrains the AI model. New training data is added during this process, improving the model's accuracy.
[0391] Input: Unidentified list, user input data
[0392] Output: Retrained AI model
[0393] Step 6:
[0394] Emotion analysis
[0395] The server uses an emotion engine to recognize the emotions of customers and users. For example, it estimates emotions from a customer's facial expressions and body movements and stores them as emotion data. Based on this, it customizes how analysis results and reports are displayed.
[0396] Input: Video data
[0397] Output: Sentiment data
[0398] Step 7:
[0399] Report generation and user interface display
[0400] The server generates reports for each observation area and store based on calculations of biodensity, customer behavior, and emotions. These reports are stored in a central database and can be accessed by users via terminals. Users can review the analysis results and reports and input information on unidentified organisms and customers.
[0401] Input: Calculation results, biological / customer data, sentiment data
[0402] Output: Biodensity report, customer behavior report, user interface display
[0403] Example of a prompt
[0404] "Please analyze the behavior of customers who spend extended periods of time in specific areas of the store (e.g., the cosmetics section), and based on the results of this analysis, including emotional data, please propose improvements."
[0405] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0406] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0407] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0408] [Second Embodiment]
[0409] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0410] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0411] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0412] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0413] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0414] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0415] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0416] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0417] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0418] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0419] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0420] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0421] This invention relates to a system that acquires video footage from multiple cameras installed at a specific observation point, analyzes the acquired footage to identify biological species and individuals, and calculates biological density. The system has a configuration centered around a server, terminals, and users, and operates as follows.
[0422] Data collection
[0423] The server periodically acquires video footage from multiple cameras installed at observation points and stores it in a central database. For example, it acquires video every hour from observation points A and B in a mountainous area.
[0424] Video analysis
[0425] The server uses an AI model to analyze the acquired video footage. This model includes object detection algorithms using deep learning (e.g., YOLO and Mask R-CNN). The server uses this to detect living organisms in the video, identify their species, and identify the characteristics of each individual. For example, it can identify three deer and two rabbits from the video of observation point A.
[0426] Database update
[0427] The server stores data on detected and identified organisms in a central database. For organisms that are not identified, the video data is stored and used later as training data. Based on the data of newly identified organisms, the AI model is retrained to improve the accuracy of the analysis. For example, a newly discovered insect could be registered in the database.
[0428] density calculation
[0429] The server calculates the biodensity for each observation area based on the analysis results. The calculation results are compiled into a report and stored in a central database. For example, at observation point A, the deer density is 1.5 individuals / m³. 2 Rabbit density is 1 rabbit / m 2 Record this result in the report.
[0430] User Interface
[0431] The terminal provides an interface for users to access the system and view analysis results and biodensity reports. Users can view results for specific observation points and input information about unidentified organisms. This input information is used to train the AI model for the next time. For example, a user might view results for observation point A and supplement the information with unidentified organisms to facilitate learning.
[0432] As described above, this system can efficiently perform a series of processes including species and individual identification within the observation area, density calculation, and user-assisted information supplementation. This makes it possible to accurately measure the number of species and populations of organisms and to improve the accuracy of the analysis year after year.
[0433] The following describes the processing flow.
[0434] Step 1: Data Collection
[0435] The server periodically collects video footage from multiple cameras installed at the observation site.
[0436] The server stores the collected video data in a temporary storage folder and then saves it as a backup in the central database.
[0437] For example, video footage is collected from observation points A and B every hour and stored in a central database.
[0438] Step 2: Video Analysis
[0439] The server reads unanalyzed video data from the central database.
[0440] The server loads a pre-trained deep learning model (e.g., YOLO or Mask R-CNN) and analyzes the video data.
[0441] The server detects organisms in the video and assigns a bounding box, species label, and confidence score to each organism. For example, it identifies three deer and two rabbits from video footage of observation point A.
[0442] Step 3: Individual Identification
[0443] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species.
[0444] The server assigns a unique ID to each individual and identifies them by considering their distinctive physical characteristics (patterns, body size, etc.). For example, if one of three deer has a different pattern, the server will use that as the basis for individual identification.
[0445] Step 4: Database Update
[0446] The server stores the species and individual identification information of the identified organisms in a database.
[0447] The server adds data of unidentified organisms to an unidentified list and saves the video footage for later use as training data. For example, it adds newly discovered insects to the unidentified list.
[0448] Step 5: Register training data and retrain the AI model
[0449] The user checks the unidentified list and enters the species name and information of the newly identified organism.
[0450] The server updates the training data based on new biological information provided by the user and prepares to retrain the AI model.
[0451] The server retrains to improve the recognition accuracy of the AI model.
[0452] Step 6: Density Calculation
[0453] The server compiles the number of organisms in each observation area based on the analysis results.
[0454] The server uses the aggregated number of organisms to determine the biodensity (number of individuals / m²) for each area. 2 Calculate the density of deer at observation point A. For example, the density of deer at observation point A is 1.5 deer / m². 2 Rabbit density is 1 rabbit / m 2 This is the result.
[0455] Step 7: Report Generation
[0456] The server generates reports for each observation area based on the calculated biodensity.
[0457] The server saves the generated reports to a central database.
[0458] Step 8: Displaying the User Interface
[0459] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user.
[0460] Users review the observation results through the interface and input information about unidentified organisms as needed. This information is then used to train the AI model for the next time.
[0461] Through the steps described above, this system can efficiently measure the number of species and population size of organisms, and its accuracy can be improved year by year.
[0462] (Example 1)
[0463] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0464] In nature observation and environmental monitoring, efficiently analyzing large amounts of video data acquired from cameras installed at specific observation points and accurately identifying species and individuals of organisms is crucial. However, conventional methods have mainly relied on manual analysis, which is labor-intensive, time-consuming, and suffers from accuracy issues. Furthermore, there has been no mechanism to effectively utilize data on unidentified organisms and continuously improve the accuracy of the analysis. This has resulted in limitations in the accuracy and reusability of observation data, making accurate calculation of biodensities difficult.
[0465] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0466] In this invention, the server includes means for acquiring video data from multiple image acquisition devices installed at observation sites, means for storing the acquired video data in a data storage device, means for inputting the stored video data into an analysis device for analysis and automatically detecting and identifying organisms in the video, means for storing the data of identified organisms in a database, means for registering the data of unidentified organisms as training data and retraining the recognition model device, and means for calculating the organism density for each observation site and generating report data. This enables efficient and highly accurate analysis of large amounts of video data, and by continuously improving the analysis accuracy, accurate calculation of organism density becomes possible.
[0467] An "image acquisition device" is a device installed at a specific observation point to acquire video data of the surrounding area.
[0468] "Video data" refers to video and still image data acquired by an image acquisition device.
[0469] A "data storage device" is a device used to store acquired video data. Examples include databases and storage servers.
[0470] An "analysis device" is a device that analyzes acquired and stored video data to automatically detect and identify biological species and individuals. It primarily uses AI models and deep learning algorithms.
[0471] "Living organisms" refers to plants, animals, and other organisms detected at the observation site.
[0472] A "database" is a system for managing and storing data on identified organisms.
[0473] "Data on organisms that could not be identified" refers to video data of organisms that the analysis device could not automatically identify.
[0474] "Training data" refers to data used to retrain an AI model.
[0475] A "recognition model device" is a device that operates and manages AI models used for identification.
[0476] "Biodensity" refers to the number of organisms per unit area at a specific observation point.
[0477] "Report data" refers to data compiled in the form of a report, which includes the results of biodensity calculations and analyses.
[0478] A "user interface" is the interface through which a user accesses a system and views analysis results and reports.
[0479] "Users" refer to individuals who access the system, view analysis results, or input information about unidentified organisms.
[0480] This invention relates to a system that acquires video data from multiple image acquisition devices installed at an observation site, analyzes the acquired video data to identify biological species and individuals, and further calculates biological density. The system has a configuration centered around a server, terminals, and users, and operates as follows.
[0481] Data collection
[0482] The server periodically acquires video data from multiple image acquisition devices (e.g., cameras) installed at observation sites and stores it in a central database. High-resolution cameras are suitable for use. For example, a configuration could be set up to acquire video data from observation sites A and B every hour. As a specific example, video taken at observation site A at 10:00 AM would be saved as "A_20231001_10.mp4".
[0483] Video analysis
[0484] The server analyzes the acquired video data using an analysis device (e.g., a high-performance GPU server). This analysis device is equipped with an object detection algorithm using deep learning (e.g., YOLO, Mask R-CNN). The server uses this AI model to detect living organisms in the video, identify each species, and analyze the characteristics of each individual. For example, one might identify three deer and two rabbits from video footage of observation point A. In this case, software such as Python and TensorFlow would be used.
[0485] Database update
[0486] The server stores data on detected and identified organisms in a database. For organisms that are not identified, their video data is saved as training data and later used to retrain the AI model. The AI model is retrained based on data of newly identified organisms to improve the accuracy of the analysis. For example, a newly discovered insect might be registered in the database, and this data could be used to improve the accuracy of the model.
[0487] density calculation
[0488] The server calculates the biodensity at each observation point based on the analysis results. These results are compiled into a report and stored in a database. For example, at observation point A, the deer density is 1.5 individuals / m³. 2 Rabbit density is 1 rabbit / m 2 This is recorded in the report. R or the Pandas library in Python are used for the calculations.
[0489] User Interface
[0490] The terminal provides an interface for users to access the system and view analysis results. This interface runs on a web browser and has the functionality to view results for a specific observation location. Users can input information about unidentified organisms, and this input information is used for the next AI model retraining. For example, a user might view the results for observation location A and supplement the information about unidentified organisms to accelerate learning. The interface is implemented using React, with Node.js and Express used for the backend.
[0491] As a concrete example, the following prompt can be input to the generating AI model:
[0492] Please describe in natural language the process of a system that acquires video footage from multiple cameras installed at observation sites, uses an AI model to detect organisms in the footage, and identifies the characteristics of species and individuals. Please explain in detail what kind of data processing and calculations are performed using specific hardware (e.g., high-resolution cameras, GPU servers) and software (e.g., Python, TensorFlow, YOLO, React, Node.js). For example, please explain how and from where data is acquired, how it is analyzed, how the analysis results are stored, how users can view the results, and how they can input information about unidentified organisms.
[0493] This invention enables efficient execution of a series of processes, including species and individual identification within an observation area, density calculation, and user-assisted information supplementation. This allows for accurate measurement of the number of species and populations, and improves the accuracy of the analysis year after year.
[0494] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0495] Step 1:
[0496] Data collection
[0497] The server acquires video data from multiple image acquisition devices installed at observation sites. Specifically, it acquires video every hour from cameras installed at observation sites A and B. The input is video data from the image acquisition devices, and the output is the acquired video data. For example, video taken at observation site A at 10:00 AM is saved as "A_20231001_10.mp4". This data is transferred to a central database via the high-speed internet.
[0498] Step 2:
[0499] Saving video data
[0500] The server stores the acquired video data in a data storage device. The server uses a MySQL database to save the video data in an appropriate format. The input is the video data acquired in step 1, and the output is the video data stored in the database. For example, "A_20231001_10.mp4" is stored in the MySQL database.
[0501] Step 3:
[0502] Video analysis
[0503] The server inputs stored video data into an analysis device (high-performance GPU server) for analysis. The server utilizes deep learning-based object detection algorithms (e.g., YOLO, Mask R-CNN). The input is video data read from a database, and the output is detected and identified biological data. For example, three deer and two rabbits are identified from video footage of observation point A. This process uses Python and TensorFlow.
[0504] Step 4:
[0505] Database update
[0506] The server stores the analyzed biological data in a database. Specifically, it stores species and individual data of detected and identified organisms. The input is the analysis result data, and the output is the biological data stored in the database. Video data of organisms that were not identified is also stored and used later as training data. For example, data for a newly discovered organism is stored as "species: unidentified, quantity: 1".
[0507] Step 5:
[0508] Model Retraining
[0509] The server retrains the AI model using data on organisms that were not identified. The server uses TensorFlow to retrain the AI model. The input is the new training data, and the output is the retrained AI model. After retraining, an evaluation test is run to confirm the improvement in accuracy.
[0510] Step 6:
[0511] density calculation
[0512] The server calculates the biodensity for each observation point based on the analysis results. The server performs the calculations using R or the Pandas library in Python. The input is the analyzed biodata, and the output is the calculated biodensity. For example, the report for observation point A might show "Deer density = 1.5 deer / m²". 2 Rabbit density = 1 rabbit / m 2 The result is as follows:
[0513] Step 7:
[0514] Report generation and storage
[0515] The server generates a report based on the calculation results and saves it to the database. The generated report is saved in PDF format. The input is the density calculation result data, and the output is the generated report. For example, it is saved as "report_A_20231001.pdf".
[0516] Step 8:
[0517] Providing a user interface
[0518] The terminal provides an interface for users to access the system and view analysis results and reports. The interface runs on a web browser and is implemented using React. Input is user access requests, and output is the display of analysis results and the ability to download reports. Users can view results for specific observation locations and input information about unidentified organisms.
[0519] Step 9:
[0520] Reflecting user input
[0521] The user inputs information about unidentified organisms through the user interface. This information is stored in a database and used for retraining the AI model the next time. The input is the information about unidentified organisms entered by the user, and the output is the updated training data.
[0522] summary
[0523] Thus, this system efficiently executes a series of processes, from data collection and analysis to database updates, density calculations, and user interface input. By using specific hardware and software, accurate identification of species and individuals, as well as density calculations, are possible, and continuous learning and accuracy improvement are achieved.
[0524] (Application Example 1)
[0525] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0526] Conventional animal observation systems focus on identifying species and individuals, but lack the ability to identify specific individuals in real time and monitor their density and intrusion status. This has resulted in inefficient security management within facilities, leading to delays in detecting suspicious individuals and issuing alarms. Furthermore, the lack of a system that simultaneously performs both animal observation and person identification has made integrated data management difficult. This invention aims to improve facility security levels by enabling real-time identification of individuals alongside animal observation.
[0527] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0528] In this invention, the server includes means for acquiring video footage from multiple cameras, means for analyzing the acquired video footage to detect and identify animals, means for storing data of detected and identified animals in a storage device, means for registering data of unidentified animals as training data and retraining an AI model, means for calculating animal density for each observation point and generating a report, means for analyzing the acquired video footage to identify specific individuals and their density in real time, means for storing data of detected and identified individuals in a database, means for calculating the density of individuals in each area based on the analysis results, means for inputting information on unidentified animals, and means for issuing alarms based on identified individuals. This makes it possible to efficiently perform both animal observation and security monitoring.
[0529] A "filming device" is a device installed at a specific observation point to capture video footage.
[0530] "Video footage" refers to a series of image data acquired from multiple recording devices.
[0531] "Identification" is the process of distinguishing individual animals or people as specific species from acquired video footage.
[0532] "Detection" is the process of finding objects in acquired video footage.
[0533] A "storage device" is hardware or software used to store data.
[0534] A "database" is a system for efficiently storing and managing structured data.
[0535] An "AI model" is an artificial intelligence program created based on machine learning algorithms and trained to perform a specific task.
[0536] "Real-time" means processing and analyzing acquired data immediately and providing the results instantly.
[0537] "Density" is an indicator that shows the number of animals or people present within a given observation area.
[0538] A "report" is a document that summarizes observational data and analysis results.
[0539] An "alarm" is a warning signal that is issued when specific conditions are met.
[0540] A "domain" refers to a specific area or place that is the subject of observation or surveillance.
[0541] A "user interface" is software that provides a means for a system and a user to interact.
[0542] "Training data" refers to the dataset used to train an AI model.
[0543] "Retraining" is the process of retraining an existing AI model using new data.
[0544] This invention relates to a system that acquires video footage from multiple cameras installed at observation points, analyzes the acquired footage to detect and identify animals and people, and uses the results for security management. The system has a configuration centered around a server, terminals, and users, and operates as follows.
[0545] Data collection
[0546] The server periodically acquires video footage from multiple cameras installed within the facility and stores it in a central database. For example, it might acquire video footage every hour from a camera installed in a specific research facility.
[0547] Video analysis
[0548] The server uses deep learning-based object detection algorithms (e.g., YOLO or Mask R-CNN) to analyze the acquired video footage. This algorithm is used to identify animals and specific individuals within the video. For example, it can identify three deer, two rabbits, and a specific person from footage from a certain camera.
[0549] Database update
[0550] The server stores data on detected and identified animals and people in a central database. For animals that are not identified, the video data is stored as training data and later used to retrain the AI model.
[0551] density calculation
[0552] The server calculates the density of animals and people in each observation area based on the analysis results. The calculation results are compiled into a report and stored in a central database. For example, at a specific observation point, the deer density is 1.5 individuals / m². 2 Rabbit density is 1 rabbit / m 2 The density of people is 0.2 people / m². 2 Record this result in the report.
[0553] Alarm and monitoring
[0554] The server has a means of issuing real-time alarms based on identified individuals. For example, if an unauthorized person enters a specific authorized area, an alarm is automatically issued and the responsible person is notified.
[0555] User Interface
[0556] The terminal provides an interface for users to access the system and view analysis results and animal and human density reports. Users can view results for specific observation locations and input information on unidentified animals. This input information will be used to train the next AI model.
[0557] Hardware and software to be used
[0558] Hardware:
[0559] Recording equipment: Surveillance cameras installed within the facility (e.g., typical network cameras)
[0560] Server: A computer that runs the database and deep learning models (e.g., typical server hardware).
[0561] software:
[0562] OpenCV: Used to capture and save camera footage.
[0563] PyTorch: A deep learning framework
[0564] SQLite: Database Administration
[0565] Flask: Building User Interfaces
[0566] Examples of processes and prompt statements
[0567] Specific example:
[0568] Cameras installed in a research facility capture video in real time to detect intrusions by specific individuals. The detection results are stored in a central database and notified to security personnel in real time.
[0569] Example of a prompt:
[0570] Analyze the surveillance camera footage within the facility to identify specific individuals and their density in real time.
[0571] For example, it checks whether a man wearing a white coat is in an authorized area.
[0572] As described above, this system can improve facility security by efficiently performing a series of processes including animal species and individual identification within the observation area, density calculation, real-time person identification and density calculation, alarm output, and user-assisted information supplementation.
[0573] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0574] Step 1:
[0575] The server acquires video footage from multiple cameras installed within the facility. The cameras periodically capture video data and send it to the server. The input is video data from the cameras, and the output is a saved video file.
[0576] Step 2:
[0577] The server analyzes the acquired video footage using deep learning-based object detection algorithms (e.g., YOLO or Mask R-CNN). Here, animals and people are detected within the video. The input is the video file acquired in step 1, and the output is a list of detected animals and people along with their location information.
[0578] Step 3:
[0579] The server stores data on detected and identified animals and people in a central database. Attribute information is also stored for each identified individual. The input is the list and location information obtained in step 2, and the output is the data stored in the database.
[0580] Step 4:
[0581] The server registers data of animals that were not identified as training data and uses it to retrain the AI model. The input is video data of animals that were not identified in step 2, and the output is the updated AI model.
[0582] Step 5:
[0583] The server calculates the density of animals and people in each observation area. Here, it calculates the number of animals and people present in a specific area and determines their density. The input is the data saved in step 3, and the output is the calculated density information.
[0584] Step 6:
[0585] The server generates a report for each observation area based on the calculation results. The report includes animal species and density, human density, and all detected data. The input is the density information obtained in step 5, and the output is the generated report.
[0586] Step 7:
[0587] The terminal provides the generated report to the user through a user interface. The user can review the results for a specific observation point and supplement information on unidentified animals. The input is the report generated in step 6, and the output is the analysis results displayed to the user.
[0588] Step 8:
[0589] The server identifies specific individuals and issues alarms in real time. For example, if an unauthorized person enters a specific authorized area, an alarm is automatically triggered. The input is data on the person detected in step 2, and the output is a record of the alarms that were triggered.
[0590] This series of processing steps enables the creation of a system that efficiently performs both animal observation and security monitoring within the facility.
[0591] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0592] This invention is a system that acquires video footage from multiple cameras installed at a specific observation point, analyzes the data to identify species and individuals, and calculates biodensity. Furthermore, it features an emotion engine that recognizes user emotions to improve the user experience. This system has a configuration centered around a server, terminals, and users, and operates as follows.
[0593] Data collection
[0594] The server periodically collects video footage from multiple cameras installed at observation points and stores it in a central database. For example, it collects video from observation points A and B every hour and stores it in the central database.
[0595] Video analysis
[0596] The server uses an AI model to analyze the acquired video footage. This model includes object detection algorithms using deep learning (e.g., YOLO and Mask R-CNN). The server uses this to detect living organisms in the video, identify their species, and identify the characteristics of each individual. For example, it can identify three deer and two rabbits from the video of observation point A.
[0597] Individual identification
[0598] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species. It assigns a unique ID to each individual and identifies them based on distinctive physical characteristics (patterns, body size, etc.). For example, if one of three deer has a different pattern, it will be identified based on that.
[0599] Database update
[0600] The server stores the species and individual identification information of identified organisms in a database. It adds data of unidentified organisms to an unidentified list and saves their images for later use as training data. For example, it adds newly discovered insects to the unidentified list.
[0601] Registration of training data and retraining of the AI model
[0602] The user reviews the unidentified list and enters the species name and information of newly identified organisms. This information is used to train the AI model for the next time. The server updates the training data based on the new biological information provided by the user and retrains the AI model. This improves the accuracy of the next analysis.
[0603] density calculation
[0604] The server aggregates the number of organisms in each observation area based on the analysis results, and calculates the biodensity (number of individuals / m³) for each area. 2 Calculate the density of deer at observation point A. For example, the density of deer at observation point A is 1.5 deer / m². 2 Rabbit density is 1 rabbit / m 2 This is the result.
[0605] Report generation
[0606] The server generates reports for each observation area based on the calculated biodensity. These generated reports are then stored in a central database.
[0607] User Interface Display
[0608] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user. Users can review results for specific observation points and input information about organisms that were not identified. This input information is used to train the AI model for the next analysis, enabling more accurate analysis.
[0609] Combination of emotional engines
[0610] The server uses an emotion engine to recognize the user's emotions. Based on this emotion data, it can customize how analysis results and reports are displayed. For example, if a user is stressed, the frequency and level of detail of information provided can be adjusted to reduce their burden. Furthermore, the emotion engine can also analyze the context and emotions of unidentified organisms entered by the user, enabling automatic completion and correction.
[0611] Specific example
[0612] For example, based on video footage acquired from observation point A, the server identifies three deer and two rabbits and assigns a unique ID to each individual. Later, when the user uses a device to view the analysis results, the emotion engine recognizes the user's interests and stress levels and customizes the displayed content. When the user discovers a new, unidentified insect and enters its information, the emotion engine analyzes the context and emotions of the input and automatically completes the information.
[0613] This system enables efficient measurement of the number of species and population size of organisms, improving analysis accuracy and optimizing the user experience.
[0614] The following describes the processing flow.
[0615] Step 1: Data Collection
[0616] The server periodically collects video footage from multiple cameras installed at the observation site.
[0617] The server stores the collected video data in a temporary storage folder and simultaneously saves it as a backup in the central database.
[0618] For example, video footage is collected from observation points A and B every hour and stored in a central database.
[0619] Step 2: Video Analysis
[0620] The server reads unanalyzed video data from the central database.
[0621] The server loads a pre-trained deep learning model (e.g., YOLO or Mask R-CNN) and analyzes the video data.
[0622] The server assigns bounding boxes, species labels, and confidence scores to organisms detected in the video. For example, it identifies three deer and two rabbits from video footage of observation point A.
[0623] Step 3: Individual Identification
[0624] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species.
[0625] The server assigns a unique ID to each individual and identifies them by considering their distinctive physical characteristics (patterns, body size, etc.). For example, if one of three deer has a different pattern, the server will use that as the basis for individual identification.
[0626] Step 4: Database Update
[0627] The server stores the species and individual identification information of the identified organisms in a database.
[0628] The server adds data of unidentified organisms to an unidentified list and saves the video footage for later use as training data. For example, it adds newly discovered insects to the unidentified list.
[0629] Step 5: Register training data and retrain the AI model
[0630] The user checks the unidentified list and enters the species name and information of the newly identified organism.
[0631] The server updates the training data based on new biological information provided by the user and prepares to retrain the AI model.
[0632] The server retrains to improve the recognition accuracy of the AI model.
[0633] Step 6: Density Calculation
[0634] The server compiles the number of organisms in each observation area based on the analysis results.
[0635] The server uses the aggregated number of organisms to determine the biodensity (number of individuals / m²) for each area. 2 Calculate the density of deer at observation point A. For example, the density of deer at observation point A is 1.5 deer / m². 2 Rabbit density is 1 rabbit / m 2 This is the result.
[0636] Step 7: Report Generation
[0637] The server generates reports for each observation area based on the calculated biodensity.
[0638] The server saves the generated reports to a central database.
[0639] Step 8: Displaying the User Interface
[0640] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user.
[0641] Users review observation results through the interface and input information about organisms that were not identified. This information is then used to train the AI model for the next time.
[0642] Step 9: Combining Emotional Engines
[0643] The server uses an emotion engine to recognize the user's emotions. For example, it analyzes sensor data from webcams and microphones to evaluate the user's facial expressions and tone of voice.
[0644] The server customizes how analysis results and reports are displayed based on emotional data. For example, if a user is experiencing stress, it adjusts the frequency and level of detail of information provided to reduce the user's burden.
[0645] Users also use the sentiment engine to analyze the context and emotions of the unidentified organisms they input. This allows for simple corrections and automatic completion, ensuring that accurate information is reflected in the database.
[0646] Through the steps described above, this system can efficiently measure the number of species and population size of organisms, improve analysis accuracy year after year, and optimize the user experience.
[0647] (Example 2)
[0648] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0649] Conventional biological observation systems can identify species when analyzing video footage from observation sites, but they lack the means to uniquely identify individual organisms. Furthermore, they fail to provide information that takes into account the user's emotional state, resulting in an unoptimized user experience. Therefore, improvements in both observation accuracy and user experience are needed.
[0650] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0651] In this invention, the server includes means for acquiring video data from multiple cameras installed at observation sites, means for analyzing the acquired video data to detect and identify organisms in the video, means for storing data of detected and identified organisms in a storage device, means for registering data of organisms that were not identified as training data and retraining an artificial intelligence model, means for calculating the organism density for each observation site and generating a report, means for uniquely identifying each individual organism using the artificial intelligence model, and means for recognizing the user's emotions and customizing the analysis results and reports based on that data. This improves the accuracy of organism observation and optimizes the user experience.
[0652] A "camera" is a device installed at an observation point to acquire video data.
[0653] "Video data" refers to video information acquired from multiple recording devices.
[0654] A "server" is a computing device that analyzes video data to detect and identify living organisms.
[0655] A "memory device" is a data storage device used to store data on detected and identified organisms.
[0656] An "artificial intelligence model" is a general term for machine learning algorithms used for detecting and identifying living organisms, as well as for uniquely identifying individuals.
[0657] "Retraining" is a learning process that uses data from organisms that were not identified to improve the accuracy of an artificial intelligence model.
[0658] "Biodensity" refers to the number of organisms per unit area at a specific observation point.
[0659] A "report" is a document that summarizes analysis results, biodensities, and other data.
[0660] "User emotions" refer to the psychological state a user experiences while using a system, such as interest or stress.
[0661] "Customization" refers to adjusting how analysis results and reports are displayed based on user sentiment data.
[0662] This invention is a system that acquires video data from multiple cameras installed at a specific observation point, analyzes that data to identify species and individuals, and calculates biodensity. Furthermore, it is characterized by its ability to improve the user experience by incorporating an emotion engine that recognizes the user's emotions. This system has a configuration centered on a server, terminals, and users, and operates as follows.
[0663] Data collection
[0664] The server periodically collects video data from multiple cameras installed at observation sites and stores it in a central database. For example, it collects video data from observation sites A and B every hour and stores it in the central database. It downloads and manages video streams using HTTP requests.
[0665] Video analysis
[0666] The server loads a pre-trained deep learning model (such as YOLO or Mask R-CNN) to analyze the acquired video data. This model is used to analyze each frame of the video, and an object detection algorithm is applied. The GPU is used to perform model inference at high speed, detecting the location and species of organisms in the video. For example, analyzing video from observation point A can detect three deer and two rabbits in each frame.
[0667] Individual identification
[0668] The server applies additional deep learning models and clustering algorithms to each detected organism to identify individuals within the same species. Each individual is assigned a unique ID and identified based on characteristics such as markings and body size. For example, if one of three deer has a different marking, it is identified as "Deer 1," "Deer 2," and "Deer 3" based on that.
[0669] Database update
[0670] The server updates the database with species and individual identification information for identified organisms. It uses SQL queries to insert the identification information into the appropriate tables. Data for unidentified organisms is added to an unidentified list, and their videos are saved separately. For example, a newly discovered insect is added to the unidentified list, and its video is saved to cloud storage.
[0671] Registration of training data and retraining of the AI model
[0672] The user reviews the unidentified list and enters the species name and other information of newly identified organisms. This operation can be performed through a form in the web application. This information is sent to the server in JSON format. The server incorporates the new organism data provided by the user into the training database and generates a new training set. The AI model is retrained using a GPU cluster to improve accuracy.
[0673] density calculation
[0674] The server aggregates the number of organisms in each observation area based on the analysis results and calculates the biodensity. It obtains area information for each observation point from the database and calculates the density by dividing the number of individuals by the area. For example, using data obtained from the scan of observation point A, it calculates that "the density of deer is 1.5 individuals / m²". 2 "The rabbit density is 1 rabbit / m²" 2 The calculation is as follows:
[0675] Report generation
[0676] The server generates reports for each observation area based on the calculated biodensity. A Python script is used to automatically generate these reports, which are then stored in a central database. These reports include density information and individual identification information for each species.
[0677] User Interface Display
[0678] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user. This interface is designed using front-end technologies such as JavaScript and React. Users can view results for specific observation points and input information about organisms that were not identified. This input information is sent to the server and used for training the next AI model.
[0679] Combination of emotional engines
[0680] The server uses an emotion engine to recognize the user's emotions. It analyzes user operation logs and input data and applies algorithms to predict the emotional state. Emotional data is reflected in report displays and alert generation. For example, if the user is stressed, the frequency of information provided may be reduced or the display mode may be switched to a simplified mode. Even when the user enters information about an unfamiliar insect, the emotion engine analyzes the context and emotions and performs automatic completion or correction.
[0681] Specific example
[0682] For example, based on video data acquired from observation point A, the server identifies three deer and two rabbits and assigns a unique ID to each individual. Subsequently, when the user uses a device to view the analysis results, the emotion engine recognizes the user's interests and stress levels and customizes the displayed content. When the user discovers a new unidentified insect and enters its information, the emotion engine analyzes the context and emotions of the input and automatically completes the information.
[0683] Example of a prompt
[0684] Analyze the following video footage and identify the species of organisms shown. Also, identify each individual organism and update the database with your findings.
[0685] For organisms on the unidentified list, please retrain the AI model based on the information entered by the user to improve the identification accuracy for the next time.
[0686] Analyze user sentiment and customize how the analysis results and reports are displayed accordingly.
[0687] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0688] Step 1:
[0689] The server periodically collects video data from multiple cameras installed at observation sites. Its input is the video stream from the observation sites, obtained via HTTP requests. Specifically, it uses a scheduling function to access the cameras every hour and download the video data. The output is the storage of the acquired video data in a central database.
[0690] Step 2:
[0691] The server loads a pre-trained deep learning model (e.g., YOLO or Mask R-CNN) to analyze the acquired video data. The input is the stored video data. Specifically, each frame of the video is divided, and an object detection algorithm is applied. The output is the location information and species identification result of organisms in each frame. Model inference is performed at high speed using a GPU, and analysis is performed efficiently.
[0692] Step 3:
[0693] The server applies additional deep learning models and clustering algorithms to the detected organisms to identify individuals within the same species. The input is the location information and species identification results of the organisms obtained in step 2. Specifically, a unique ID is assigned to each individual, and features such as patterns and body size are analyzed. The output is the uniquely identified individual information. This makes it possible to make specific identifications such as "Deer 1," "Deer 2," and "Deer 3."
[0694] Step 4:
[0695] The server updates the database with the species and individual identification information of the identified organisms. The input is the unique individual information obtained in step 3. Specifically, it inserts this information into the appropriate table in the database using an SQL query. The output is the updated database information. Data of organisms that were not identified is added to the unidentified list, and their video data is saved to cloud storage.
[0696] Step 5:
[0697] The user reviews the unidentified list and enters the species name and other information of newly identified organisms. This input is provided through a web application form. This information is sent to the server in JSON format. The server incorporates this new data into the training database and generates a new training set. The output is an updated training data set and a new training set to be used for the next training session.
[0698] Step 6:
[0699] The server aggregates the number of organisms in each observation area based on the analysis results and calculates the biodensity. The input is the database information updated in step 4. Specifically, it obtains area information for each observation point from the database and calculates the density by dividing the number of individuals by the area. The output is the biodensity information for each observation area. For example, from the data for observation point A, it might say "the deer density is 1.5 individuals / m²". 2 "The rabbit density is 1 rabbit / m²" 2 The calculation is as follows:
[0700] Step 7:
[0701] The server generates reports for each observation area based on biodensity. The input is the biodensity information obtained in step 6. A Python script is used to automatically generate the reports and save them to the central database. The output is the generated report. This report includes density information and individual identification information for each species.
[0702] Step 8:
[0703] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user. Input consists of reports and analysis results obtained from the server. The web page is designed using front-end technologies such as JavaScript and React. Output consists of analysis results and reports displayed to the user. The user can review results for specific observation areas and add information about organisms that were not identified.
[0704] Step 9:
[0705] The server uses an emotion engine to recognize the user's emotions. Input is the user's operation logs and input data. An algorithm that predicts emotional states is used to detect stress, interest, etc. The output is the recognized emotion information. Based on this, the frequency of information provision and the display method are automatically customized. Even when the user inputs information about an unknown insect, the emotion engine analyzes the context and emotions and automatically completes or corrects the information.
[0706] (Application Example 2)
[0707] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0708] There is a need to accurately understand customer behavior and emotions in stores and use that information to improve product placement and service quality, but current technology makes it difficult to fully achieve this. Furthermore, there is a lack of efficient methods for calculating and identifying biological densities. Therefore, more advanced data analysis technologies are needed to improve the efficiency of store operations and enhance the customer experience.
[0709] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0710] In this invention, the server includes means for acquiring video footage from multiple cameras installed at observation points, means for analyzing the acquired video footage to detect and identify organisms in the footage, means for storing data of the detected and identified organisms in a database, means for registering data of organisms that were not identified as training data and retraining an AI model, means for calculating the organism density for each observation point and generating a report, means for filming and analyzing customer behavior from multiple cameras installed in the store, means for analyzing the video footage to identify customers and assign a unique ID, and means for identifying emotional data using an emotion engine that analyzes customer emotions. This makes it possible to analyze customer behavior and emotions in detail, optimize store operations based on this analysis, and improve the customer experience.
[0711] An "observation point" is a specific area or location where the system acquires video footage and uses it for analysis.
[0712] A "camera" is a device used to capture video footage and collect that data.
[0713] A "server" is a central processing unit used to analyze collected video footage, store data, and retrain AI models.
[0714] "Video footage" is a media format composed of a series of still images that records the situation at an observation point or inside a store in real time.
[0715] "Living organisms" refers to species and individuals of animals, plants, insects, and other organisms present at the observation site.
[0716] "Identification" refers to the act of identifying a specific species, individual, or customer from analyzed video footage.
[0717] A "database" is an information system for systematically storing collected and analyzed data.
[0718] An "AI model" is an algorithm that uses deep learning or machine learning to perform data analysis.
[0719] "Retraining" is the process of improving the accuracy of an AI model using new training data.
[0720] "Biological density" is an index calculated by determining the number of organisms in a specific observation location per unit area.
[0721] A "report" is a quantitative and qualitative report generated based on the results of an analysis.
[0722] A "customer" is a person who engages in activities within a store.
[0723] An "emotion engine" is an analytical device and software that estimates emotions from a customer's facial expressions and body movements.
[0724] A "unique ID" is identification information that assigns a unique identifier to each identified organism or customer.
[0725] This invention is a system that acquires video footage from multiple cameras installed at specific observation points, analyzes the data to identify biological species and individuals, and calculates biodensity. It also aims to improve the user experience by incorporating an emotion engine that recognizes user emotions. Furthermore, by applying this system to physical stores and analyzing customer behavior and emotions, it aims to optimize store operations and improve the quality of customer service.
[0726] 1. Data Collection
[0727] The server periodically collects video footage from multiple cameras installed at observation points and stores it in a central database. For example, it collects video from observation points A and B every hour and stores it in the central database. It also periodically acquires video data from cameras installed inside the store. This allows for continuous monitoring of customer behavior.
[0728] 2. Video Analysis
[0729] The server uses the following AI model to analyze the acquired video footage.
[0730] Object detection algorithms such as YOLO (You Only Look Once) or Mask R-CNN
[0731] These models are used to detect organisms in video footage and identify their species and individual characteristics. Additionally, in in-store video footage, customers are identified, and each customer is assigned a unique ID.
[0732] 3. Individual Identification
[0733] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species or each customer. For example, it can identify individuals based on unique physical characteristics or body movements.
[0734] 4. Database Update
[0735] The server stores the species and individual identification information of identified organisms in a database. Unidentified data is added to an unidentified list, and its video is saved for later use as training data. The same applies to customers, storing customer behavior data and sentiment data in the database.
[0736] 5. Registering training data and retraining the AI model
[0737] Users can review the unidentified list and input the species name and information of newly identified organisms. Based on this, the server retrains the AI model. Similarly, for customer data, the model is retrained based on customer behavior patterns and sentiment data.
[0738] 6. Emotion analysis
[0739] The server uses an emotion engine to recognize the user's emotions. Based on this emotion data, it customizes how analysis results and reports are displayed. For example, if a user is feeling stressed in a store, the frequency and level of detail of information provided will be adjusted. It is also possible to analyze customer emotion data and provide appropriate services.
[0740] 7. Report generation and user interface display
[0741] The server generates reports for each observation area and store based on calculations of biodensity and customer behavior. These reports are stored in a central database and accessible to users via terminals. Users can review the analysis results and reports and input information about unidentified organisms and customers.
[0742] Specific example
[0743] For example, based on video footage acquired from observation point A, the server identifies three deer and two rabbits and assigns a unique ID to each individual. When the user then reviews the analysis results, the emotion engine recognizes the user's interests and stress levels and customizes the displayed content. Furthermore, when the user discovers a new unidentified insect and inputs its information, the emotion engine analyzes the context and emotions of the input and automatically completes the information.
[0744] Example of a prompt
[0745] "Please analyze the behavior of customers who spend extended periods of time in specific areas of the store (e.g., the cosmetics section), and based on the results of this analysis, including emotional data, please propose improvements."
[0746] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0747] Step 1:
[0748] Data collection
[0749] The server periodically collects video footage from multiple cameras installed at observation points and within the store. For example, the server collects video from observation points A and B every hour and stores it in a central database. It also acquires video data from cameras inside the store and stores it in the central database as well.
[0750] Input: Camera footage
[0751] Output: Video data stored in the central database
[0752] Step 2:
[0753] Video analysis
[0754] The server uses an AI model to analyze the acquired video footage. Specifically, it employs object detection algorithms such as YOLO and Mask R-CNN. This allows it to detect living beings and customers in the video and assign a unique ID to each individual or customer.
[0755] Input: Video data
[0756] Output: Identified individual / customer data (with unique ID)
[0757] Step 3:
[0758] Individual identification
[0759] The server further analyzes the detailed characteristics of identified organisms and customers using deep learning models and clustering algorithms. It uniquely identifies each individual or customer based on their physical characteristics and behavioral patterns.
[0760] Input: Identified individual / customer data
[0761] Output: Detailed data on individual customers (feature-based unique IDs)
[0762] Step 4:
[0763] Database update
[0764] The server stores data on identified organisms and customers in a database. Data that is not identified is added to an unidentified list, and this video is saved and used as training data later.
[0765] Input: Individual / customer detailed data, unidentified data
[0766] Output: Updated database
[0767] Step 5:
[0768] Registration of training data and retraining of the AI model
[0769] The user reviews the unidentified list and inputs information about newly identified organisms or customers. Based on this information, the server retrains the AI model. New training data is added during this process, improving the model's accuracy.
[0770] Input: Unidentified list, user input data
[0771] Output: Retrained AI model
[0772] Step 6:
[0773] Emotion analysis
[0774] The server uses an emotion engine to recognize the emotions of customers and users. For example, it estimates emotions from a customer's facial expressions and body movements and stores them as emotion data. Based on this, it customizes how analysis results and reports are displayed.
[0775] Input: Video data
[0776] Output: Sentiment data
[0777] Step 7:
[0778] Report generation and user interface display
[0779] The server generates reports for each observation area and store based on calculations of biodensity, customer behavior, and emotions. These reports are stored in a central database and can be accessed by users via terminals. Users can review the analysis results and reports and input information about unidentified organisms and customers.
[0780] Input: Calculation results, biological / customer data, sentiment data
[0781] Output: Biodensity report, customer behavior report, user interface display
[0782] Example of a prompt
[0783] "Please analyze the behavior of customers who spend extended periods of time in specific areas of the store (e.g., the cosmetics section), and based on the results of this analysis, including emotional data, please propose improvements."
[0784] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0785] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0786] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0787] [Third Embodiment]
[0788] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0789] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0790] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0791] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0792] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0793] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0794] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0795] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0796] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0797] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0798] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0799] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0800] This invention relates to a system that acquires video footage from multiple cameras installed at a specific observation point, analyzes the acquired footage to identify biological species and individuals, and calculates biological density. The system has a configuration centered around a server, terminals, and users, and operates as follows.
[0801] Data collection
[0802] The server periodically acquires video footage from multiple cameras installed at observation points and stores it in a central database. For example, it acquires video every hour from observation points A and B in a mountainous area.
[0803] Video analysis
[0804] The server uses an AI model to analyze the acquired video footage. This model includes object detection algorithms using deep learning (e.g., YOLO and Mask R-CNN). The server uses this to detect living organisms in the video, identify their species, and identify the characteristics of each individual. For example, it can identify three deer and two rabbits from the video of observation point A.
[0805] Database update
[0806] The server stores data on detected and identified organisms in a central database. For organisms that are not identified, the video data is stored and used later as training data. Based on the data of newly identified organisms, the AI model is retrained to improve the accuracy of the analysis. For example, a newly discovered insect could be registered in the database.
[0807] density calculation
[0808] The server calculates the biodensity for each observation area based on the analysis results. The calculation results are compiled into a report and stored in a central database. For example, at observation point A, the deer density is 1.5 individuals / m³. 2 Rabbit density is 1 rabbit / m 2 Record this result in the report.
[0809] User Interface
[0810] The terminal provides an interface for users to access the system and view analysis results and biodensity reports. Users can view results for specific observation points and input information about unidentified organisms. This input information is used to train the AI model for the next time. For example, a user might view results for observation point A and supplement the information with unidentified organisms to facilitate learning.
[0811] As described above, this system can efficiently perform a series of processes including species and individual identification within the observation area, density calculation, and user-assisted information supplementation. This makes it possible to accurately measure the number of species and populations of organisms and to improve the accuracy of the analysis year after year.
[0812] The following describes the processing flow.
[0813] Step 1: Data Collection
[0814] The server periodically collects video footage from multiple cameras installed at the observation site.
[0815] The server stores the collected video data in a temporary storage folder and then saves it as a backup in the central database.
[0816] For example, video footage is collected from observation points A and B every hour and stored in a central database.
[0817] Step 2: Video Analysis
[0818] The server reads unanalyzed video data from the central database.
[0819] The server loads a pre-trained deep learning model (e.g., YOLO or Mask R-CNN) and analyzes the video data.
[0820] The server detects organisms in the video and assigns a bounding box, species label, and confidence score to each organism. For example, it identifies three deer and two rabbits from video footage of observation point A.
[0821] Step 3: Individual Identification
[0822] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species.
[0823] The server assigns a unique ID to each individual and identifies them by considering their distinctive physical characteristics (patterns, body size, etc.). For example, if one of three deer has a different pattern, the server will use that as the basis for individual identification.
[0824] Step 4: Database Update
[0825] The server stores the species and individual identification information of the identified organisms in a database.
[0826] The server adds data of unidentified organisms to an unidentified list and saves the video footage for later use as training data. For example, it adds newly discovered insects to the unidentified list.
[0827] Step 5: Register training data and retrain the AI model
[0828] The user checks the unidentified list and enters the species name and information of the newly identified organism.
[0829] The server updates the training data based on new biological information provided by the user and prepares to retrain the AI model.
[0830] The server retrains to improve the recognition accuracy of the AI model.
[0831] Step 6: Density Calculation
[0832] The server compiles the number of organisms in each observation area based on the analysis results.
[0833] The server uses the aggregated number of organisms to determine the biodensity (number of individuals / m²) for each area. 2 Calculate the density of deer at observation point A. For example, the density of deer at observation point A is 1.5 deer / m². 2 Rabbit density is 1 rabbit / m 2 This is the result.
[0834] Step 7: Report Generation
[0835] The server generates reports for each observation area based on the calculated biodensity.
[0836] The server saves the generated reports to a central database.
[0837] Step 8: Displaying the User Interface
[0838] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user.
[0839] Users review the observation results through the interface and input information about unidentified organisms as needed. This information is then used to train the AI model for the next time.
[0840] Through the steps described above, this system can efficiently measure the number of species and population size of organisms, and its accuracy can be improved year by year.
[0841] (Example 1)
[0842] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0843] In nature observation and environmental monitoring, efficiently analyzing large amounts of video data acquired from cameras installed at specific observation points and accurately identifying species and individuals of organisms is crucial. However, conventional methods have mainly relied on manual analysis, which is labor-intensive, time-consuming, and suffers from accuracy issues. Furthermore, there has been no mechanism to effectively utilize data on unidentified organisms and continuously improve the accuracy of the analysis. This has resulted in limitations in the accuracy and reusability of observation data, making accurate calculation of biodensities difficult.
[0844] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0845] In this invention, the server includes means for acquiring video data from multiple image acquisition devices installed at observation sites, means for storing the acquired video data in a data storage device, means for inputting the stored video data into an analysis device for analysis and automatically detecting and identifying organisms in the video, means for storing the data of identified organisms in a database, means for registering the data of unidentified organisms as training data and retraining the recognition model device, and means for calculating the organism density for each observation site and generating report data. This enables efficient and highly accurate analysis of large amounts of video data, and by continuously improving the analysis accuracy, accurate calculation of organism density becomes possible.
[0846] An "image acquisition device" is a device installed at a specific observation point to acquire video data of the surrounding area.
[0847] "Video data" refers to video and still image data acquired by an image acquisition device.
[0848] A "data storage device" is a device used to store acquired video data. Examples include databases and storage servers.
[0849] An "analysis device" is a device that analyzes acquired and stored video data to automatically detect and identify biological species and individuals. It primarily uses AI models and deep learning algorithms.
[0850] "Living organisms" refers to plants, animals, and other organisms detected at the observation site.
[0851] A "database" is a system for managing and storing data on identified organisms.
[0852] "Data on organisms that could not be identified" refers to video data of organisms that the analysis device could not automatically identify.
[0853] "Training data" refers to data used to retrain an AI model.
[0854] A "recognition model device" is a device that operates and manages AI models used for identification.
[0855] "Biodensity" refers to the number of organisms per unit area at a specific observation point.
[0856] "Report data" refers to data compiled in the form of a report, which includes the results of biodensity calculations and analyses.
[0857] A "user interface" is the interface through which a user accesses a system and views analysis results and reports.
[0858] "Users" refer to individuals who access the system, view analysis results, or input information about unidentified organisms.
[0859] This invention relates to a system that acquires video data from multiple image acquisition devices installed at an observation site, analyzes the acquired video data to identify biological species and individuals, and further calculates biological density. The system has a configuration centered around a server, terminals, and users, and operates as follows.
[0860] Data collection
[0861] The server periodically acquires video data from multiple image acquisition devices (e.g., cameras) installed at observation sites and stores it in a central database. High-resolution cameras are suitable for use. For example, a configuration could be set up to acquire video data from observation sites A and B every hour. As a specific example, video taken at observation site A at 10:00 AM would be saved as "A_20231001_10.mp4".
[0862] Video analysis
[0863] The server analyzes the acquired video data using an analysis device (e.g., a high-performance GPU server). This analysis device is equipped with an object detection algorithm using deep learning (e.g., YOLO, Mask R-CNN). The server uses this AI model to detect living organisms in the video, identify each species, and analyze the characteristics of each individual. For example, one might identify three deer and two rabbits from video footage of observation point A. In this case, software such as Python and TensorFlow would be used.
[0864] Database update
[0865] The server stores data on detected and identified organisms in a database. For organisms that are not identified, their video data is saved as training data and later used to retrain the AI model. The AI model is retrained based on data of newly identified organisms to improve the accuracy of the analysis. For example, a newly discovered insect might be registered in the database, and this data could be used to improve the accuracy of the model.
[0866] density calculation
[0867] The server calculates the biodensity at each observation point based on the analysis results. These results are compiled into a report and stored in a database. For example, at observation point A, the deer density is 1.5 individuals / m³. 2 Rabbit density is 1 rabbit / m 2 This is recorded in the report. R or the Pandas library in Python are used for the calculations.
[0868] User Interface
[0869] The terminal provides an interface for users to access the system and view analysis results. This interface runs on a web browser and has the functionality to view results for a specific observation location. Users can input information about unidentified organisms, and this input information is used for the next AI model retraining. For example, a user might view the results for observation location A and supplement the information about unidentified organisms to accelerate learning. The interface is implemented using React, with Node.js and Express used for the backend.
[0870] As a concrete example, the following prompt can be input to the generating AI model:
[0871] Please describe in natural language the process of a system that acquires video footage from multiple cameras installed at observation sites, uses an AI model to detect organisms in the footage, and identifies the characteristics of species and individuals. Please explain in detail what kind of data processing and calculations are performed using specific hardware (e.g., high-resolution cameras, GPU servers) and software (e.g., Python, TensorFlow, YOLO, React, Node.js). For example, please explain how and from where data is acquired, how it is analyzed, how the analysis results are stored, how users can view the results, and how they can input information about unidentified organisms.
[0872] This invention enables efficient execution of a series of processes, including species and individual identification within an observation area, density calculation, and user-assisted information supplementation. This allows for accurate measurement of the number of species and populations, and improves the accuracy of the analysis year after year.
[0873] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0874] Step 1:
[0875] Data collection
[0876] The server acquires video data from multiple image acquisition devices installed at observation sites. Specifically, it acquires video every hour from cameras installed at observation sites A and B. The input is video data from the image acquisition devices, and the output is the acquired video data. For example, video taken at observation site A at 10:00 AM is saved as "A_20231001_10.mp4". This data is transferred to a central database via the high-speed internet.
[0877] Step 2:
[0878] Saving video data
[0879] The server stores the acquired video data in a data storage device. The server uses a MySQL database to save the video data in an appropriate format. The input is the video data acquired in step 1, and the output is the video data stored in the database. For example, "A_20231001_10.mp4" is stored in the MySQL database.
[0880] Step 3:
[0881] Video analysis
[0882] The server inputs stored video data into an analysis device (high-performance GPU server) for analysis. The server utilizes deep learning-based object detection algorithms (e.g., YOLO, Mask R-CNN). The input is video data read from a database, and the output is detected and identified biological data. For example, three deer and two rabbits are identified from video footage of observation point A. This process uses Python and TensorFlow.
[0883] Step 4:
[0884] Database update
[0885] The server stores the analyzed biological data in a database. Specifically, it stores species and individual data of detected and identified organisms. The input is the analysis result data, and the output is the biological data stored in the database. Video data of organisms that were not identified is also stored and used later as training data. For example, data for a newly discovered organism is stored as "species: unidentified, quantity: 1".
[0886] Step 5:
[0887] Model Retraining
[0888] The server retrains the AI model using data on organisms that were not identified. The server uses TensorFlow to retrain the AI model. The input is the new training data, and the output is the retrained AI model. After retraining, an evaluation test is run to confirm the improvement in accuracy.
[0889] Step 6:
[0890] density calculation
[0891] The server calculates the biodensity for each observation point based on the analysis results. The server performs the calculations using R or the Pandas library in Python. The input is the analyzed biodata, and the output is the calculated biodensity. For example, the report for observation point A might show "Deer density = 1.5 deer / m²". 2 Rabbit density = 1 rabbit / m 2 The result is as follows:
[0892] Step 7:
[0893] Report generation and storage
[0894] The server generates a report based on the calculation results and saves it to the database. The generated report is saved in PDF format. The input is the density calculation result data, and the output is the generated report. For example, it is saved as "report_A_20231001.pdf".
[0895] Step 8:
[0896] Providing a user interface
[0897] The terminal provides an interface for users to access the system and view analysis results and reports. The interface runs on a web browser and is implemented using React. Input is user access requests, and output is the display of analysis results and the ability to download reports. Users can view results for specific observation locations and input information about unidentified organisms.
[0898] Step 9:
[0899] Reflecting user input
[0900] The user inputs information about unidentified organisms through the user interface. This information is stored in a database and used for retraining the AI model the next time. The input is the information about unidentified organisms entered by the user, and the output is the updated training data.
[0901] summary
[0902] Thus, this system efficiently executes a series of processes, from data collection and analysis to database updates, density calculations, and user interface input. By using specific hardware and software, accurate identification of species and individuals, as well as density calculations, are possible, and continuous learning and accuracy improvement are achieved.
[0903] (Application Example 1)
[0904] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0905] Conventional animal observation systems focus on identifying species and individuals, but lack the ability to identify specific individuals in real time and monitor their density and intrusion status. This has resulted in inefficient security management within facilities, leading to delays in detecting suspicious individuals and issuing alarms. Furthermore, the lack of a system that simultaneously performs both animal observation and person identification has made integrated data management difficult. This invention aims to improve facility security levels by enabling real-time identification of individuals alongside animal observation.
[0906] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0907] In this invention, the server includes means for acquiring video footage from multiple cameras, means for analyzing the acquired video footage to detect and identify animals, means for storing data of detected and identified animals in a storage device, means for registering data of unidentified animals as training data and retraining an AI model, means for calculating animal density for each observation point and generating a report, means for analyzing the acquired video footage to identify specific individuals and their density in real time, means for storing data of detected and identified individuals in a database, means for calculating the density of individuals in each area based on the analysis results, means for inputting information on unidentified animals, and means for issuing alarms based on identified individuals. This makes it possible to efficiently perform both animal observation and security monitoring.
[0908] A "filming device" is a device installed at a specific observation point to capture video footage.
[0909] "Video footage" refers to a series of image data acquired from multiple recording devices.
[0910] "Identification" is the process of distinguishing individual animals or people as specific species from acquired video footage.
[0911] "Detection" is the process of finding objects in acquired video footage.
[0912] A "storage device" is hardware or software used to store data.
[0913] A "database" is a system for efficiently storing and managing structured data.
[0914] An "AI model" is an artificial intelligence program created based on machine learning algorithms and trained to perform a specific task.
[0915] "Real-time" means processing and analyzing acquired data immediately and providing the results instantly.
[0916] "Density" is an indicator that shows the number of animals or people present within a given observation area.
[0917] A "report" is a document that summarizes observational data and analysis results.
[0918] An "alarm" is a warning signal that is issued when specific conditions are met.
[0919] A "domain" refers to a specific area or place that is the subject of observation or surveillance.
[0920] A "user interface" is software that provides a means for a system and a user to interact.
[0921] "Training data" refers to the dataset used to train an AI model.
[0922] "Retraining" is the process of retraining an existing AI model using new data.
[0923] This invention relates to a system that acquires video footage from multiple cameras installed at observation points, analyzes the acquired footage to detect and identify animals and people, and uses the results for security management. The system has a configuration centered around a server, terminals, and users, and operates as follows.
[0924] Data collection
[0925] The server periodically acquires video footage from multiple cameras installed within the facility and stores it in a central database. For example, it might acquire video footage every hour from a camera installed in a specific research facility.
[0926] Video analysis
[0927] The server uses deep learning-based object detection algorithms (e.g., YOLO or Mask R-CNN) to analyze the acquired video footage. This algorithm is used to identify animals and specific individuals within the video. For example, it can identify three deer, two rabbits, and a specific person from footage from a certain camera.
[0928] Database update
[0929] The server stores data on detected and identified animals and people in a central database. For animals that are not identified, the video data is stored as training data and later used to retrain the AI model.
[0930] density calculation
[0931] The server calculates the density of animals and people in each observation area based on the analysis results. The calculation results are compiled into a report and stored in a central database. For example, at a specific observation point, the deer density is 1.5 individuals / m². 2 Rabbit density is 1 rabbit / m 2 The density of people is 0.2 people / m². 2 Record this result in the report.
[0932] Alarm and monitoring
[0933] The server has a means of issuing real-time alarms based on identified individuals. For example, if an unauthorized person enters a specific authorized area, an alarm is automatically issued and the responsible person is notified.
[0934] User Interface
[0935] The terminal provides an interface for users to access the system and view analysis results and animal and human density reports. Users can view results for specific observation locations and input information on unidentified animals. This input information will be used to train the next AI model.
[0936] Hardware and software to be used
[0937] Hardware:
[0938] Recording equipment: Surveillance cameras installed within the facility (e.g., typical network cameras)
[0939] Server: A computer that runs the database and deep learning models (e.g., typical server hardware).
[0940] software:
[0941] OpenCV: Used to capture and save camera footage.
[0942] PyTorch: A deep learning framework
[0943] SQLite: Database Administration
[0944] Flask: Building User Interfaces
[0945] Examples of processes and prompt statements
[0946] Specific example:
[0947] Cameras installed in a research facility capture video in real time to detect intrusions by specific individuals. The detection results are stored in a central database and notified to security personnel in real time.
[0948] Example of a prompt:
[0949] Analyze the surveillance camera footage within the facility to identify specific individuals and their density in real time.
[0950] For example, it checks whether a man wearing a white coat is in an authorized area.
[0951] As described above, this system can improve facility security by efficiently performing a series of processes including animal species and individual identification within the observation area, density calculation, real-time person identification and density calculation, alarm output, and user-assisted information supplementation.
[0952] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0953] Step 1:
[0954] The server acquires video footage from multiple cameras installed within the facility. The cameras periodically capture video data and send it to the server. The input is video data from the cameras, and the output is a saved video file.
[0955] Step 2:
[0956] The server analyzes the acquired video footage using deep learning-based object detection algorithms (e.g., YOLO or Mask R-CNN). Here, animals and people are detected within the video. The input is the video file acquired in step 1, and the output is a list of detected animals and people along with their location information.
[0957] Step 3:
[0958] The server stores data on detected and identified animals and people in a central database. Attribute information is also stored for each identified individual. The input is the list and location information obtained in step 2, and the output is the data stored in the database.
[0959] Step 4:
[0960] The server registers data of animals that were not identified as training data and uses it to retrain the AI model. The input is video data of animals that were not identified in step 2, and the output is the updated AI model.
[0961] Step 5:
[0962] The server calculates the density of animals and people in each observation area. Here, it calculates the number of animals and people present in a specific area and determines their density. The input is the data saved in step 3, and the output is the calculated density information.
[0963] Step 6:
[0964] The server generates a report for each observation area based on the calculation results. The report includes animal species and density, human density, and all detected data. The input is the density information obtained in step 5, and the output is the generated report.
[0965] Step 7:
[0966] The terminal provides the generated report to the user through a user interface. The user can review the results for a specific observation point and supplement information on unidentified animals. The input is the report generated in step 6, and the output is the analysis results displayed to the user.
[0967] Step 8:
[0968] The server identifies specific individuals and issues alarms in real time. For example, if an unauthorized person enters a specific authorized area, an alarm is automatically triggered. The input is data on the person detected in step 2, and the output is a record of the alarms that were triggered.
[0969] This series of processing steps enables the creation of a system that efficiently performs both animal observation and security monitoring within the facility.
[0970] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0971] This invention is a system that acquires video footage from multiple cameras installed at a specific observation point, analyzes the data to identify species and individuals, and calculates biodensity. Furthermore, it features an emotion engine that recognizes user emotions to improve the user experience. This system has a configuration centered around a server, terminals, and users, and operates as follows.
[0972] Data collection
[0973] The server periodically collects video footage from multiple cameras installed at observation points and stores it in a central database. For example, it collects video from observation points A and B every hour and stores it in the central database.
[0974] Video analysis
[0975] The server uses an AI model to analyze the acquired video footage. This model includes object detection algorithms using deep learning (e.g., YOLO and Mask R-CNN). The server uses this to detect living organisms in the video, identify their species, and identify the characteristics of each individual. For example, it can identify three deer and two rabbits from the video of observation point A.
[0976] Individual identification
[0977] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species. It assigns a unique ID to each individual and identifies them based on distinctive physical characteristics (patterns, body size, etc.). For example, if one of three deer has a different pattern, it will be identified based on that.
[0978] Database update
[0979] The server stores the species and individual identification information of identified organisms in a database. It adds data of unidentified organisms to an unidentified list and saves their images for later use as training data. For example, it adds newly discovered insects to the unidentified list.
[0980] Registration of training data and retraining of the AI model
[0981] The user reviews the unidentified list and enters the species name and information of newly identified organisms. This information is used to train the AI model for the next time. The server updates the training data based on the new biological information provided by the user and retrains the AI model. This improves the accuracy of the next analysis.
[0982] density calculation
[0983] The server aggregates the number of organisms in each observation area based on the analysis results, and calculates the biodensity (number of individuals / m³) for each area. 2 Calculate the density of deer at observation point A. For example, the density of deer at observation point A is 1.5 deer / m². 2 Rabbit density is 1 rabbit / m 2 This is the result.
[0984] Report generation
[0985] The server generates reports for each observation area based on the calculated biodensity. These generated reports are then stored in a central database.
[0986] User Interface Display
[0987] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user. Users can review results for specific observation points and input information about organisms that were not identified. This input information is used to train the AI model for the next analysis, enabling more accurate analysis.
[0988] Combination of emotional engines
[0989] The server uses an emotion engine to recognize the user's emotions. Based on this emotion data, it can customize how analysis results and reports are displayed. For example, if a user is stressed, the frequency and level of detail of information provided can be adjusted to reduce their burden. Furthermore, the emotion engine can also analyze the context and emotions of unidentified organisms entered by the user, enabling automatic completion and correction.
[0990] Specific example
[0991] For example, based on video footage acquired from observation point A, the server identifies three deer and two rabbits and assigns a unique ID to each individual. Later, when the user uses a device to view the analysis results, the emotion engine recognizes the user's interests and stress levels and customizes the displayed content. When the user discovers a new, unidentified insect and enters its information, the emotion engine analyzes the context and emotions of the input and automatically completes the information.
[0992] This system enables efficient measurement of the number of species and population size of organisms, improving analysis accuracy and optimizing the user experience.
[0993] The following describes the processing flow.
[0994] Step 1: Data Collection
[0995] The server periodically collects video footage from multiple cameras installed at the observation site.
[0996] The server stores the collected video data in a temporary storage folder and simultaneously saves it as a backup in the central database.
[0997] For example, video footage is collected from observation points A and B every hour and stored in a central database.
[0998] Step 2: Video Analysis
[0999] The server reads unanalyzed video data from the central database.
[1000] The server loads a pre-trained deep learning model (e.g., YOLO or Mask R-CNN) and analyzes the video data.
[1001] The server assigns bounding boxes, species labels, and confidence scores to organisms detected in the video. For example, it identifies three deer and two rabbits from video footage of observation point A.
[1002] Step 3: Individual Identification
[1003] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species.
[1004] The server assigns a unique ID to each individual and identifies them by considering their distinctive physical characteristics (patterns, body size, etc.). For example, if one of three deer has a different pattern, the server will use that as the basis for individual identification.
[1005] Step 4: Database Update
[1006] The server stores the species and individual identification information of the identified organisms in a database.
[1007] The server adds data of unidentified organisms to an unidentified list and saves the video footage for later use as training data. For example, it adds newly discovered insects to the unidentified list.
[1008] Step 5: Register training data and retrain the AI model
[1009] The user checks the unidentified list and enters the species name and information of the newly identified organism.
[1010] The server updates the training data based on new biological information provided by the user and prepares to retrain the AI model.
[1011] The server retrains to improve the recognition accuracy of the AI model.
[1012] Step 6: Density Calculation
[1013] The server compiles the number of organisms in each observation area based on the analysis results.
[1014] The server uses the aggregated number of organisms to determine the biodensity (number of individuals / m²) for each area. 2 Calculate the density of deer at observation point A. For example, the density of deer at observation point A is 1.5 deer / m². 2 Rabbit density is 1 rabbit / m 2 This is the result.
[1015] Step 7: Report Generation
[1016] The server generates reports for each observation area based on the calculated biodensity.
[1017] The server saves the generated reports to a central database.
[1018] Step 8: Displaying the User Interface
[1019] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user.
[1020] Users review observation results through the interface and input information about organisms that were not identified. This information is then used to train the AI model for the next time.
[1021] Step 9: Combining Emotional Engines
[1022] The server uses an emotion engine to recognize the user's emotions. For example, it analyzes sensor data from webcams and microphones to evaluate the user's facial expressions and tone of voice.
[1023] The server customizes how analysis results and reports are displayed based on emotional data. For example, if a user is experiencing stress, it adjusts the frequency and level of detail of information provided to reduce the user's burden.
[1024] Users also use the sentiment engine to analyze the context and emotions of the unidentified organisms they input. This allows for simple corrections and automatic completion, ensuring that accurate information is reflected in the database.
[1025] Through the steps described above, this system can efficiently measure the number of species and population size of organisms, improve analysis accuracy year after year, and optimize the user experience.
[1026] (Example 2)
[1027] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1028] Conventional biological observation systems can identify species when analyzing video footage from observation sites, but they lack the means to uniquely identify individual organisms. Furthermore, they fail to provide information that takes into account the user's emotional state, resulting in an unoptimized user experience. Therefore, improvements in both observation accuracy and user experience are needed.
[1029] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1030] In this invention, the server includes means for acquiring video data from multiple cameras installed at observation sites, means for analyzing the acquired video data to detect and identify organisms in the video, means for storing data of detected and identified organisms in a storage device, means for registering data of organisms that were not identified as training data and retraining an artificial intelligence model, means for calculating the organism density for each observation site and generating a report, means for uniquely identifying each individual organism using the artificial intelligence model, and means for recognizing the user's emotions and customizing the analysis results and reports based on that data. This improves the accuracy of organism observation and optimizes the user experience.
[1031] A "camera" is a device installed at an observation point to acquire video data.
[1032] "Video data" refers to video information acquired from multiple recording devices.
[1033] A "server" is a computing device that analyzes video data to detect and identify living organisms.
[1034] A "memory device" is a data storage device used to store data on detected and identified organisms.
[1035] An "artificial intelligence model" is a general term for machine learning algorithms used for detecting and identifying living organisms, as well as for uniquely identifying individuals.
[1036] "Retraining" is a learning process that uses data from organisms that were not identified to improve the accuracy of an artificial intelligence model.
[1037] "Biodensity" refers to the number of organisms per unit area at a specific observation point.
[1038] A "report" is a document that summarizes analysis results, biodensities, and other data.
[1039] "User emotions" refer to the psychological state a user experiences while using a system, such as interest or stress.
[1040] "Customization" refers to adjusting how analysis results and reports are displayed based on user sentiment data.
[1041] This invention is a system that acquires video data from multiple cameras installed at a specific observation point, analyzes that data to identify species and individuals, and calculates biodensity. Furthermore, it is characterized by its ability to improve the user experience by incorporating an emotion engine that recognizes the user's emotions. This system has a configuration centered on a server, terminals, and users, and operates as follows.
[1042] Data collection
[1043] The server periodically collects video data from multiple cameras installed at observation sites and stores it in a central database. For example, it collects video data from observation sites A and B every hour and stores it in the central database. It downloads and manages video streams using HTTP requests.
[1044] Video analysis
[1045] The server loads a pre-trained deep learning model (such as YOLO or Mask R-CNN) to analyze the acquired video data. This model is used to analyze each frame of the video, and an object detection algorithm is applied. The GPU is used to perform model inference at high speed, detecting the location and species of organisms in the video. For example, analyzing video from observation point A can detect three deer and two rabbits in each frame.
[1046] Individual identification
[1047] The server applies additional deep learning models and clustering algorithms to each detected organism to identify individuals within the same species. Each individual is assigned a unique ID and identified based on characteristics such as markings and body size. For example, if one of three deer has a different marking, it is identified as "Deer 1," "Deer 2," and "Deer 3" based on that.
[1048] Database update
[1049] The server updates the database with species and individual identification information for identified organisms. It uses SQL queries to insert the identification information into the appropriate tables. Data for unidentified organisms is added to an unidentified list, and their videos are saved separately. For example, a newly discovered insect is added to the unidentified list, and its video is saved to cloud storage.
[1050] Registration of training data and retraining of the AI model
[1051] The user reviews the unidentified list and enters the species name and other information of newly identified organisms. This operation can be performed through a form in the web application. This information is sent to the server in JSON format. The server incorporates the new organism data provided by the user into the training database and generates a new training set. The AI model is retrained using a GPU cluster to improve accuracy.
[1052] density calculation
[1053] The server aggregates the number of organisms in each observation area based on the analysis results and calculates the biodensity. It obtains area information for each observation point from the database and calculates the density by dividing the number of individuals by the area. For example, using data obtained from the scan of observation point A, it calculates that "the density of deer is 1.5 individuals / m²". 2 "The rabbit density is 1 rabbit / m²" 2 The calculation is as follows:
[1054] Report generation
[1055] The server generates reports for each observation area based on the calculated biodensity. A Python script is used to automatically generate these reports, which are then stored in a central database. These reports include density information and individual identification information for each species.
[1056] User Interface Display
[1057] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user. This interface is designed using front-end technologies such as JavaScript and React. Users can view results for specific observation points and input information about organisms that were not identified. This input information is sent to the server and used for training the next AI model.
[1058] Combination of emotional engines
[1059] The server uses an emotion engine to recognize the user's emotions. It analyzes user operation logs and input data and applies algorithms to predict the emotional state. Emotional data is reflected in report displays and alert generation. For example, if the user is stressed, the frequency of information provided may be reduced or the display mode may be switched to a simplified mode. Even when the user enters information about an unfamiliar insect, the emotion engine analyzes the context and emotions and performs automatic completion or correction.
[1060] Specific example
[1061] For example, based on video data acquired from observation point A, the server identifies three deer and two rabbits and assigns a unique ID to each individual. Subsequently, when the user uses a device to view the analysis results, the emotion engine recognizes the user's interests and stress levels and customizes the displayed content. When the user discovers a new unidentified insect and enters its information, the emotion engine analyzes the context and emotions of the input and automatically completes the information.
[1062] Example of a prompt
[1063] Analyze the following video footage and identify the species of organisms shown. Also, identify each individual organism and update the database with your findings.
[1064] For organisms on the unidentified list, please retrain the AI model based on the information entered by the user to improve the identification accuracy for the next time.
[1065] Analyze user sentiment and customize how the analysis results and reports are displayed accordingly.
[1066] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1067] Step 1:
[1068] The server periodically collects video data from multiple cameras installed at observation sites. Its input is the video stream from the observation sites, obtained via HTTP requests. Specifically, it uses a scheduling function to access the cameras every hour and download the video data. The output is the storage of the acquired video data in a central database.
[1069] Step 2:
[1070] The server loads a pre-trained deep learning model (e.g., YOLO or Mask R-CNN) to analyze the acquired video data. The input is the stored video data. Specifically, each frame of the video is divided, and an object detection algorithm is applied. The output is the location information and species identification result of organisms in each frame. Model inference is performed at high speed using a GPU, and analysis is performed efficiently.
[1071] Step 3:
[1072] The server applies additional deep learning models and clustering algorithms to the detected organisms to identify individuals within the same species. The input is the location information and species identification results of the organisms obtained in step 2. Specifically, a unique ID is assigned to each individual, and features such as patterns and body size are analyzed. The output is the uniquely identified individual information. This makes it possible to make specific identifications such as "Deer 1," "Deer 2," and "Deer 3."
[1073] Step 4:
[1074] The server updates the database with the species and individual identification information of the identified organisms. The input is the unique individual information obtained in step 3. Specifically, it inserts this information into the appropriate table in the database using an SQL query. The output is the updated database information. Data of organisms that were not identified is added to the unidentified list, and their video data is saved to cloud storage.
[1075] Step 5:
[1076] The user reviews the unidentified list and enters the species name and other information of newly identified organisms. This input is provided through a web application form. This information is sent to the server in JSON format. The server incorporates this new data into the training database and generates a new training set. The output is an updated training data set and a new training set to be used for the next training session.
[1077] Step 6:
[1078] The server aggregates the number of organisms in each observation area based on the analysis results and calculates the biodensity. The input is the database information updated in step 4. Specifically, it obtains area information for each observation point from the database and calculates the density by dividing the number of individuals by the area. The output is the biodensity information for each observation area. For example, from the data for observation point A, it might say "the deer density is 1.5 individuals / m²". 2 "The rabbit density is 1 rabbit / m²" 2 The calculation is as follows:
[1079] Step 7:
[1080] The server generates reports for each observation area based on biodensity. The input is the biodensity information obtained in step 6. A Python script is used to automatically generate the reports and save them to the central database. The output is the generated report. This report includes density information and individual identification information for each species.
[1081] Step 8:
[1082] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user. Input consists of reports and analysis results obtained from the server. The web page is designed using front-end technologies such as JavaScript and React. Output consists of analysis results and reports displayed to the user. The user can review results for specific observation areas and add information about organisms that were not identified.
[1083] Step 9:
[1084] The server uses an emotion engine to recognize the user's emotions. Input is the user's operation logs and input data. An algorithm that predicts emotional states is used to detect stress, interest, etc. The output is the recognized emotion information. Based on this, the frequency of information provision and the display method are automatically customized. Even when the user inputs information about an unknown insect, the emotion engine analyzes the context and emotions and automatically completes or corrects the information.
[1085] (Application Example 2)
[1086] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1087] There is a need to accurately understand customer behavior and emotions in stores and use that information to improve product placement and service quality, but current technology makes it difficult to fully achieve this. Furthermore, there is a lack of efficient methods for calculating and identifying biological densities. Therefore, more advanced data analysis technologies are needed to improve the efficiency of store operations and enhance the customer experience.
[1088] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1089] In this invention, the server includes means for acquiring video footage from multiple cameras installed at observation points, means for analyzing the acquired video footage to detect and identify organisms in the footage, means for storing data of the detected and identified organisms in a database, means for registering data of organisms that were not identified as training data and retraining an AI model, means for calculating the organism density for each observation point and generating a report, means for filming and analyzing customer behavior from multiple cameras installed in the store, means for analyzing the video footage to identify customers and assign a unique ID, and means for identifying emotional data using an emotion engine that analyzes customer emotions. This makes it possible to analyze customer behavior and emotions in detail, optimize store operations based on this analysis, and improve the customer experience.
[1090] An "observation point" is a specific area or location where the system acquires video footage and uses it for analysis.
[1091] A "camera" is a device used to capture video footage and collect that data.
[1092] A "server" is a central processing unit used to analyze collected video footage, store data, and retrain AI models.
[1093] "Video footage" is a media format composed of a series of still images that records the situation at an observation point or inside a store in real time.
[1094] "Living organisms" refers to species and individuals of animals, plants, insects, and other organisms present at the observation site.
[1095] "Identification" refers to the act of identifying a specific species, individual, or customer from analyzed video footage.
[1096] A "database" is an information system for systematically storing collected and analyzed data.
[1097] An "AI model" is an algorithm that uses deep learning or machine learning to perform data analysis.
[1098] "Retraining" is the process of improving the accuracy of an AI model using new training data.
[1099] "Biological density" is an index calculated by determining the number of organisms in a specific observation location per unit area.
[1100] A "report" is a quantitative and qualitative report generated based on the results of an analysis.
[1101] A "customer" is a person who engages in activities within a store.
[1102] An "emotion engine" is an analytical device and software that estimates emotions from a customer's facial expressions and body movements.
[1103] A "unique ID" is identification information that assigns a unique identifier to each identified organism or customer.
[1104] This invention is a system that acquires video footage from multiple cameras installed at specific observation points, analyzes the data to identify biological species and individuals, and calculates biodensity. It also aims to improve the user experience by incorporating an emotion engine that recognizes user emotions. Furthermore, by applying this system to physical stores and analyzing customer behavior and emotions, it aims to optimize store operations and improve the quality of customer service.
[1105] 1. Data Collection
[1106] The server periodically collects video footage from multiple cameras installed at observation points and stores it in a central database. For example, it collects video from observation points A and B every hour and stores it in the central database. It also periodically acquires video data from cameras installed inside the store. This allows for continuous monitoring of customer behavior.
[1107] 2. Video Analysis
[1108] The server uses the following AI model to analyze the acquired video footage.
[1109] Object detection algorithms such as YOLO (You Only Look Once) or Mask R-CNN
[1110] These models are used to detect organisms in video footage and identify their species and individual characteristics. Additionally, in in-store video footage, customers are identified, and each customer is assigned a unique ID.
[1111] 3. Individual Identification
[1112] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species or each customer. For example, it can identify individuals based on unique physical characteristics or body movements.
[1113] 4. Database Update
[1114] The server stores the species and individual identification information of identified organisms in a database. Unidentified data is added to an unidentified list, and its video is saved for later use as training data. The same applies to customers, storing customer behavior data and sentiment data in the database.
[1115] 5. Registering training data and retraining the AI model
[1116] Users can review the unidentified list and input the species name and information of newly identified organisms. Based on this, the server retrains the AI model. Similarly, for customer data, the model is retrained based on customer behavior patterns and sentiment data.
[1117] 6. Emotion analysis
[1118] The server uses an emotion engine to recognize the user's emotions. Based on this emotion data, it customizes how analysis results and reports are displayed. For example, if a user is feeling stressed in a store, the frequency and level of detail of information provided will be adjusted. It is also possible to analyze customer emotion data and provide appropriate services.
[1119] 7. Report generation and user interface display
[1120] The server generates reports for each observation area and store based on calculations of biodensity and customer behavior. These reports are stored in a central database and accessible to users via terminals. Users can review the analysis results and reports and input information about unidentified organisms and customers.
[1121] Specific example
[1122] For example, based on video footage acquired from observation point A, the server identifies three deer and two rabbits and assigns a unique ID to each individual. When the user then reviews the analysis results, the emotion engine recognizes the user's interests and stress levels and customizes the displayed content. Furthermore, when the user discovers a new unidentified insect and inputs its information, the emotion engine analyzes the context and emotions of the input and automatically completes the information.
[1123] Example of a prompt
[1124] "Please analyze the behavior of customers who spend extended periods of time in specific areas of the store (e.g., the cosmetics section), and based on the results of this analysis, including emotional data, please propose improvements."
[1125] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1126] Step 1:
[1127] Data collection
[1128] The server periodically collects video footage from multiple cameras installed at observation points and within the store. For example, the server collects video from observation points A and B every hour and stores it in a central database. It also acquires video data from cameras inside the store and stores it in the central database as well.
[1129] Input: Camera footage
[1130] Output: Video data stored in the central database
[1131] Step 2:
[1132] Video analysis
[1133] The server uses an AI model to analyze the acquired video footage. Specifically, it employs object detection algorithms such as YOLO and Mask R-CNN. This allows it to detect living beings and customers in the video and assign a unique ID to each individual or customer.
[1134] Input: Video data
[1135] Output: Identified individual / customer data (with unique ID)
[1136] Step 3:
[1137] Individual identification
[1138] The server further analyzes the detailed characteristics of identified organisms and customers using deep learning models and clustering algorithms. It uniquely identifies each individual or customer based on their physical characteristics and behavioral patterns.
[1139] Input: Identified individual / customer data
[1140] Output: Detailed data on individual customers (feature-based unique IDs)
[1141] Step 4:
[1142] Database update
[1143] The server stores data on identified organisms and customers in a database. Data that is not identified is added to an unidentified list, and this video is saved and used as training data later.
[1144] Input: Individual / customer detailed data, unidentified data
[1145] Output: Updated database
[1146] Step 5:
[1147] Registration of training data and retraining of the AI model
[1148] The user reviews the unidentified list and inputs information about newly identified organisms or customers. Based on this information, the server retrains the AI model. New training data is added during this process, improving the model's accuracy.
[1149] Input: Unidentified list, user input data
[1150] Output: Retrained AI model
[1151] Step 6:
[1152] Emotion analysis
[1153] The server uses an emotion engine to recognize the emotions of customers and users. For example, it estimates emotions from a customer's facial expressions and body movements and stores them as emotion data. Based on this, it customizes how analysis results and reports are displayed.
[1154] Input: Video data
[1155] Output: Sentiment data
[1156] Step 7:
[1157] Report generation and user interface display
[1158] The server generates reports for each observation area and store based on calculations of biodensity, customer behavior, and emotions. These reports are stored in a central database and can be accessed by users via terminals. Users can review the analysis results and reports and input information about unidentified organisms and customers.
[1159] Input: Calculation results, biological / customer data, sentiment data
[1160] Output: Biodensity report, customer behavior report, user interface display
[1161] Example of a prompt
[1162] "Please analyze the behavior of customers who spend extended periods of time in specific areas of the store (e.g., the cosmetics section), and based on the results of this analysis, including emotional data, please propose improvements."
[1163] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1164] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1165] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1166] [Fourth Embodiment]
[1167] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1168] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1169] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1170] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1171] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1172] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1173] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1174] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1175] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1176] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1177] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1178] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1179] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1180] This invention relates to a system that acquires video footage from multiple cameras installed at a specific observation point, analyzes the acquired footage to identify biological species and individuals, and calculates biological density. The system has a configuration centered around a server, terminals, and users, and operates as follows.
[1181] Data collection
[1182] The server periodically acquires video footage from multiple cameras installed at observation points and stores it in a central database. For example, it acquires video every hour from observation points A and B in a mountainous area.
[1183] Video analysis
[1184] The server uses an AI model to analyze the acquired video footage. This model includes object detection algorithms using deep learning (e.g., YOLO and Mask R-CNN). The server uses this to detect living organisms in the video, identify their species, and identify the characteristics of each individual. For example, it can identify three deer and two rabbits from the video of observation point A.
[1185] Database update
[1186] The server stores data on detected and identified organisms in a central database. For organisms that are not identified, the video data is stored and used later as training data. Based on the data of newly identified organisms, the AI model is retrained to improve the accuracy of the analysis. For example, a newly discovered insect could be registered in the database.
[1187] density calculation
[1188] The server calculates the biodensity for each observation area based on the analysis results. The calculation results are compiled into a report and stored in a central database. For example, at observation point A, the deer density is 1.5 individuals / m³. 2 Rabbit density is 1 rabbit / m 2 Record this result in the report.
[1189] User Interface
[1190] The terminal provides an interface for users to access the system and view analysis results and biodensity reports. Users can view results for specific observation points and input information about unidentified organisms. This input information is used to train the AI model for the next time. For example, a user might view results for observation point A and supplement the information with unidentified organisms to facilitate learning.
[1191] As described above, this system can efficiently perform a series of processes including species and individual identification within the observation area, density calculation, and user-assisted information supplementation. This makes it possible to accurately measure the number of species and populations of organisms and to improve the accuracy of the analysis year after year.
[1192] The following describes the processing flow.
[1193] Step 1: Data Collection
[1194] The server periodically collects video footage from multiple cameras installed at the observation site.
[1195] The server stores the collected video data in a temporary storage folder and then saves it as a backup in the central database.
[1196] For example, video footage is collected from observation points A and B every hour and stored in a central database.
[1197] Step 2: Video Analysis
[1198] The server reads unanalyzed video data from the central database.
[1199] The server loads a pre-trained deep learning model (e.g., YOLO or Mask R-CNN) and analyzes the video data.
[1200] The server detects organisms in the video and assigns a bounding box, species label, and confidence score to each organism. For example, it identifies three deer and two rabbits from video footage of observation point A.
[1201] Step 3: Individual Identification
[1202] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species.
[1203] The server assigns a unique ID to each individual and identifies them by considering their distinctive physical characteristics (patterns, body size, etc.). For example, if one of three deer has a different pattern, the server will use that as the basis for individual identification.
[1204] Step 4: Database Update
[1205] The server stores the species and individual identification information of the identified organisms in a database.
[1206] The server adds data of unidentified organisms to an unidentified list and saves the video footage for later use as training data. For example, it adds newly discovered insects to the unidentified list.
[1207] Step 5: Register training data and retrain the AI model
[1208] The user checks the unidentified list and enters the species name and information of the newly identified organism.
[1209] The server updates the training data based on new biological information provided by the user and prepares to retrain the AI model.
[1210] The server retrains to improve the recognition accuracy of the AI model.
[1211] Step 6: Density Calculation
[1212] The server compiles the number of organisms in each observation area based on the analysis results.
[1213] The server uses the aggregated number of organisms to determine the biodensity (number of individuals / m²) for each area. 2 Calculate the density of deer at observation point A. For example, the density of deer at observation point A is 1.5 deer / m². 2 Rabbit density is 1 rabbit / m 2 This is the result.
[1214] Step 7: Report Generation
[1215] The server generates reports for each observation area based on the calculated biodensity.
[1216] The server saves the generated reports to a central database.
[1217] Step 8: Displaying the User Interface
[1218] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user.
[1219] Users review the observation results through the interface and input information about unidentified organisms as needed. This information is then used to train the AI model for the next time.
[1220] Through the steps described above, this system can efficiently measure the number of species and population size of organisms, and its accuracy can be improved year by year.
[1221] (Example 1)
[1222] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1223] In nature observation and environmental monitoring, efficiently analyzing large amounts of video data acquired from cameras installed at specific observation points and accurately identifying species and individuals of organisms is crucial. However, conventional methods have mainly relied on manual analysis, which is labor-intensive, time-consuming, and suffers from accuracy issues. Furthermore, there has been no mechanism to effectively utilize data on unidentified organisms and continuously improve the accuracy of the analysis. This has resulted in limitations in the accuracy and reusability of observation data, making accurate calculation of biodensities difficult.
[1224] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1225] In this invention, the server includes means for acquiring video data from multiple image acquisition devices installed at observation sites, means for storing the acquired video data in a data storage device, means for inputting the stored video data into an analysis device for analysis and automatically detecting and identifying organisms in the video, means for storing the data of identified organisms in a database, means for registering the data of unidentified organisms as training data and retraining the recognition model device, and means for calculating the organism density for each observation site and generating report data. This enables efficient and highly accurate analysis of large amounts of video data, and by continuously improving the analysis accuracy, accurate calculation of organism density becomes possible.
[1226] An "image acquisition device" is a device installed at a specific observation point to acquire video data of the surrounding area.
[1227] "Video data" refers to video and still image data acquired by an image acquisition device.
[1228] A "data storage device" is a device used to store acquired video data. Examples include databases and storage servers.
[1229] An "analysis device" is a device that analyzes acquired and stored video data to automatically detect and identify biological species and individuals. It primarily uses AI models and deep learning algorithms.
[1230] "Living organisms" refers to plants, animals, and other organisms detected at the observation site.
[1231] A "database" is a system for managing and storing data on identified organisms.
[1232] "Data on organisms that could not be identified" refers to video data of organisms that the analysis device could not automatically identify.
[1233] "Training data" refers to data used to retrain an AI model.
[1234] A "recognition model device" is a device that operates and manages AI models used for identification.
[1235] "Biodensity" refers to the number of organisms per unit area at a specific observation point.
[1236] "Report data" refers to data compiled in the form of a report, which includes the results of biodensity calculations and analyses.
[1237] A "user interface" is the interface through which a user accesses a system and views analysis results and reports.
[1238] "Users" refer to individuals who access the system, view analysis results, or input information about unidentified organisms.
[1239] This invention relates to a system that acquires video data from multiple image acquisition devices installed at an observation site, analyzes the acquired video data to identify biological species and individuals, and further calculates biological density. The system has a configuration centered around a server, terminals, and users, and operates as follows.
[1240] Data collection
[1241] The server periodically acquires video data from multiple image acquisition devices (e.g., cameras) installed at observation sites and stores it in a central database. High-resolution cameras are suitable for use. For example, a configuration could be set up to acquire video data from observation sites A and B every hour. As a specific example, video taken at observation site A at 10:00 AM would be saved as "A_20231001_10.mp4".
[1242] Video analysis
[1243] The server analyzes the acquired video data using an analysis device (e.g., a high-performance GPU server). This analysis device is equipped with an object detection algorithm using deep learning (e.g., YOLO, Mask R-CNN). The server uses this AI model to detect living organisms in the video, identify each species, and analyze the characteristics of each individual. For example, one might identify three deer and two rabbits from video footage of observation point A. In this case, software such as Python and TensorFlow would be used.
[1244] Database update
[1245] The server stores data on detected and identified organisms in a database. For organisms that are not identified, their video data is saved as training data and later used to retrain the AI model. The AI model is retrained based on data of newly identified organisms to improve the accuracy of the analysis. For example, a newly discovered insect might be registered in the database, and this data could be used to improve the accuracy of the model.
[1246] density calculation
[1247] The server calculates the biodensity at each observation point based on the analysis results. These results are compiled into a report and stored in a database. For example, at observation point A, the deer density is 1.5 individuals / m³. 2 Rabbit density is 1 rabbit / m 2 This is recorded in the report. R or the Pandas library in Python are used for the calculations.
[1248] User Interface
[1249] The terminal provides an interface for users to access the system and view analysis results. This interface runs on a web browser and has the functionality to view results for a specific observation location. Users can input information about unidentified organisms, and this input information is used for the next AI model retraining. For example, a user might view the results for observation location A and supplement the information about unidentified organisms to accelerate learning. The interface is implemented using React, with Node.js and Express used for the backend.
[1250] As a concrete example, the following prompt can be input to the generating AI model:
[1251] Please describe in natural language the process of a system that acquires video footage from multiple cameras installed at observation sites, uses an AI model to detect organisms in the footage, and identifies the characteristics of species and individuals. Please explain in detail what kind of data processing and calculations are performed using specific hardware (e.g., high-resolution cameras, GPU servers) and software (e.g., Python, TensorFlow, YOLO, React, Node.js). For example, please explain how and from where data is acquired, how it is analyzed, how the analysis results are stored, how users can view the results, and how they can input information about unidentified organisms.
[1252] This invention enables efficient execution of a series of processes, including species and individual identification within an observation area, density calculation, and user-assisted information supplementation. This allows for accurate measurement of the number of species and populations, and improves the accuracy of the analysis year after year.
[1253] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1254] Step 1:
[1255] Data collection
[1256] The server acquires video data from multiple image acquisition devices installed at observation sites. Specifically, it acquires video every hour from cameras installed at observation sites A and B. The input is video data from the image acquisition devices, and the output is the acquired video data. For example, video taken at observation site A at 10:00 AM is saved as "A_20231001_10.mp4". This data is transferred to a central database via the high-speed internet.
[1257] Step 2:
[1258] Saving video data
[1259] The server stores the acquired video data in a data storage device. The server uses a MySQL database to save the video data in an appropriate format. The input is the video data acquired in step 1, and the output is the video data stored in the database. For example, "A_20231001_10.mp4" is stored in the MySQL database.
[1260] Step 3:
[1261] Video analysis
[1262] The server inputs stored video data into an analysis device (high-performance GPU server) for analysis. The server utilizes deep learning-based object detection algorithms (e.g., YOLO, Mask R-CNN). The input is video data read from a database, and the output is detected and identified biological data. For example, three deer and two rabbits are identified from video footage of observation point A. This process uses Python and TensorFlow.
[1263] Step 4:
[1264] Database update
[1265] The server stores the analyzed biological data in a database. Specifically, it stores species and individual data of detected and identified organisms. The input is the analysis result data, and the output is the biological data stored in the database. Video data of organisms that were not identified is also stored and used later as training data. For example, data for a newly discovered organism is stored as "species: unidentified, quantity: 1".
[1266] Step 5:
[1267] Model Retraining
[1268] The server retrains the AI model using data on organisms that were not identified. The server uses TensorFlow to retrain the AI model. The input is the new training data, and the output is the retrained AI model. After retraining, an evaluation test is run to confirm the improvement in accuracy.
[1269] Step 6:
[1270] density calculation
[1271] The server calculates the biodensity for each observation point based on the analysis results. The server performs the calculations using R or the Pandas library in Python. The input is the analyzed biodata, and the output is the calculated biodensity. For example, the report for observation point A might show "Deer density = 1.5 deer / m²". 2 Rabbit density = 1 rabbit / m 2 The result is as follows:
[1272] Step 7:
[1273] Report generation and storage
[1274] The server generates a report based on the calculation results and saves it to the database. The generated report is saved in PDF format. The input is the density calculation result data, and the output is the generated report. For example, it is saved as "report_A_20231001.pdf".
[1275] Step 8:
[1276] Providing a user interface
[1277] The terminal provides an interface for users to access the system and view analysis results and reports. The interface runs on a web browser and is implemented using React. Input is user access requests, and output is the display of analysis results and the ability to download reports. Users can view results for specific observation locations and input information about unidentified organisms.
[1278] Step 9:
[1279] Reflecting user input
[1280] The user inputs information about unidentified organisms through the user interface. This information is stored in a database and used for retraining the AI model the next time. The input is the information about unidentified organisms entered by the user, and the output is the updated training data.
[1281] summary
[1282] Thus, this system efficiently executes a series of processes, from data collection and analysis to database updates, density calculations, and user interface input. By using specific hardware and software, accurate identification of species and individuals, as well as density calculations, are possible, and continuous learning and accuracy improvement are achieved.
[1283] (Application Example 1)
[1284] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1285] Conventional animal observation systems focus on identifying species and individuals, but lack the ability to identify specific individuals in real time and monitor their density and intrusion status. This has resulted in inefficient security management within facilities, leading to delays in detecting suspicious individuals and issuing alarms. Furthermore, the lack of a system that simultaneously performs both animal observation and person identification has made integrated data management difficult. This invention aims to improve facility security levels by enabling real-time identification of individuals alongside animal observation.
[1286] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1287] In this invention, the server includes means for acquiring video footage from multiple cameras, means for analyzing the acquired video footage to detect and identify animals, means for storing data of detected and identified animals in a storage device, means for registering data of unidentified animals as training data and retraining an AI model, means for calculating animal density for each observation point and generating a report, means for analyzing the acquired video footage to identify specific individuals and their density in real time, means for storing data of detected and identified individuals in a database, means for calculating the density of individuals in each area based on the analysis results, means for inputting information on unidentified animals, and means for issuing alarms based on identified individuals. This makes it possible to efficiently perform both animal observation and security monitoring.
[1288] A "filming device" is a device installed at a specific observation point to capture video footage.
[1289] "Video footage" refers to a series of image data acquired from multiple recording devices.
[1290] "Identification" is the process of distinguishing individual animals or people as specific species from acquired video footage.
[1291] "Detection" is the process of finding objects in acquired video footage.
[1292] A "storage device" is hardware or software used to store data.
[1293] A "database" is a system for efficiently storing and managing structured data.
[1294] An "AI model" is an artificial intelligence program created based on machine learning algorithms and trained to perform a specific task.
[1295] "Real-time" means processing and analyzing acquired data immediately and providing the results instantly.
[1296] "Density" is an indicator that shows the number of animals or people present within a given observation area.
[1297] A "report" is a document that summarizes observational data and analysis results.
[1298] An "alarm" is a warning signal that is issued when specific conditions are met.
[1299] A "domain" refers to a specific area or place that is the subject of observation or surveillance.
[1300] A "user interface" is software that provides a means for a system and a user to interact.
[1301] "Training data" refers to the dataset used to train an AI model.
[1302] "Retraining" is the process of retraining an existing AI model using new data.
[1303] This invention relates to a system that acquires video footage from multiple cameras installed at observation points, analyzes the acquired footage to detect and identify animals and people, and uses the results for security management. The system has a configuration centered around a server, terminals, and users, and operates as follows.
[1304] Data collection
[1305] The server periodically acquires video footage from multiple cameras installed within the facility and stores it in a central database. For example, it might acquire video footage every hour from a camera installed in a specific research facility.
[1306] Video analysis
[1307] The server uses deep learning-based object detection algorithms (e.g., YOLO or Mask R-CNN) to analyze the acquired video footage. This algorithm is used to identify animals and specific individuals within the video. For example, it can identify three deer, two rabbits, and a specific person from footage from a certain camera.
[1308] Database update
[1309] The server stores data on detected and identified animals and people in a central database. For animals that are not identified, the video data is stored as training data and later used to retrain the AI model.
[1310] density calculation
[1311] The server calculates the density of animals and people in each observation area based on the analysis results. The calculation results are compiled into a report and stored in a central database. For example, at a specific observation point, the deer density is 1.5 individuals / m². 2 Rabbit density is 1 rabbit / m 2 The density of people is 0.2 people / m². 2 Record this result in the report.
[1312] Alarm and monitoring
[1313] The server has a means of issuing real-time alarms based on identified individuals. For example, if an unauthorized person enters a specific authorized area, an alarm is automatically issued and the responsible person is notified.
[1314] User Interface
[1315] The terminal provides an interface for users to access the system and view analysis results and animal and human density reports. Users can view results for specific observation locations and input information on unidentified animals. This input information will be used to train the next AI model.
[1316] Hardware and software to be used
[1317] Hardware:
[1318] Recording equipment: Surveillance cameras installed within the facility (e.g., typical network cameras)
[1319] Server: A computer that runs the database and deep learning models (e.g., typical server hardware).
[1320] software:
[1321] OpenCV: Used to capture and save camera footage.
[1322] PyTorch: A deep learning framework
[1323] SQLite: Database Administration
[1324] Flask: Building User Interfaces
[1325] Examples of processes and prompt statements
[1326] Specific example:
[1327] Cameras installed in a research facility capture video in real time to detect intrusions by specific individuals. The detection results are stored in a central database and notified to security personnel in real time.
[1328] Example of a prompt:
[1329] Analyze the surveillance camera footage within the facility to identify specific individuals and their density in real time.
[1330] For example, it checks whether a man wearing a white coat is in an authorized area.
[1331] As described above, this system can improve facility security by efficiently performing a series of processes including animal species and individual identification within the observation area, density calculation, real-time person identification and density calculation, alarm output, and user-assisted information supplementation.
[1332] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1333] Step 1:
[1334] The server acquires video footage from multiple cameras installed within the facility. The cameras periodically capture video data and send it to the server. The input is video data from the cameras, and the output is a saved video file.
[1335] Step 2:
[1336] The server analyzes the acquired video footage using deep learning-based object detection algorithms (e.g., YOLO or Mask R-CNN). Here, animals and people are detected within the video. The input is the video file acquired in step 1, and the output is a list of detected animals and people along with their location information.
[1337] Step 3:
[1338] The server stores data on detected and identified animals and people in a central database. Attribute information is also stored for each identified individual. The input is the list and location information obtained in step 2, and the output is the data stored in the database.
[1339] Step 4:
[1340] The server registers data of animals that were not identified as training data and uses it to retrain the AI model. The input is video data of animals that were not identified in step 2, and the output is the updated AI model.
[1341] Step 5:
[1342] The server calculates the density of animals and people in each observation area. Here, it calculates the number of animals and people present in a specific area and determines their density. The input is the data saved in step 3, and the output is the calculated density information.
[1343] Step 6:
[1344] The server generates a report for each observation area based on the calculation results. The report includes animal species and density, human density, and all detected data. The input is the density information obtained in step 5, and the output is the generated report.
[1345] Step 7:
[1346] The terminal provides the generated report to the user through a user interface. The user can review the results for a specific observation point and supplement information on unidentified animals. The input is the report generated in step 6, and the output is the analysis results displayed to the user.
[1347] Step 8:
[1348] The server identifies specific individuals and issues alarms in real time. For example, if an unauthorized person enters a specific authorized area, an alarm is automatically triggered. The input is data on the person detected in step 2, and the output is a record of the alarms that were triggered.
[1349] This series of processing steps enables the creation of a system that efficiently performs both animal observation and security monitoring within the facility.
[1350] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1351] This invention is a system that acquires video footage from multiple cameras installed at a specific observation point, analyzes the data to identify species and individuals, and calculates biodensity. Furthermore, it features an emotion engine that recognizes user emotions to improve the user experience. This system has a configuration centered around a server, terminals, and users, and operates as follows.
[1352] Data collection
[1353] The server periodically collects video footage from multiple cameras installed at observation points and stores it in a central database. For example, it collects video from observation points A and B every hour and stores it in the central database.
[1354] Video analysis
[1355] The server uses an AI model to analyze the acquired video footage. This model includes object detection algorithms using deep learning (e.g., YOLO and Mask R-CNN). The server uses this to detect living organisms in the video, identify their species, and identify the characteristics of each individual. For example, it can identify three deer and two rabbits from the video of observation point A.
[1356] Individual identification
[1357] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species. It assigns a unique ID to each individual and identifies them based on distinctive physical characteristics (patterns, body size, etc.). For example, if one of three deer has a different pattern, it will be identified based on that.
[1358] Database update
[1359] The server stores the species and individual identification information of identified organisms in a database. It adds data of unidentified organisms to an unidentified list and saves their images for later use as training data. For example, it adds newly discovered insects to the unidentified list.
[1360] Registration of training data and retraining of the AI model
[1361] The user reviews the unidentified list and enters the species name and information of newly identified organisms. This information is used to train the AI model for the next time. The server updates the training data based on the new biological information provided by the user and retrains the AI model. This improves the accuracy of the next analysis.
[1362] density calculation
[1363] The server aggregates the number of organisms in each observation area based on the analysis results, and calculates the biodensity (number of individuals / m³) for each area. 2 Calculate the density of deer at observation point A. For example, the density of deer at observation point A is 1.5 deer / m². 2 Rabbit density is 1 rabbit / m 2 This is the result.
[1364] Report generation
[1365] The server generates reports for each observation area based on the calculated biodensity. These generated reports are then stored in a central database.
[1366] User Interface Display
[1367] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user. Users can review results for specific observation points and input information about organisms that were not identified. This input information is used to train the AI model for the next analysis, enabling more accurate analysis.
[1368] Combination of emotional engines
[1369] The server uses an emotion engine to recognize the user's emotions. Based on this emotion data, it can customize how analysis results and reports are displayed. For example, if a user is stressed, the frequency and level of detail of information provided can be adjusted to reduce their burden. Furthermore, the emotion engine can also analyze the context and emotions of unidentified organisms entered by the user, enabling automatic completion and correction.
[1370] Specific example
[1371] For example, based on video footage acquired from observation point A, the server identifies three deer and two rabbits and assigns a unique ID to each individual. Later, when the user uses a device to view the analysis results, the emotion engine recognizes the user's interests and stress levels and customizes the displayed content. When the user discovers a new, unidentified insect and enters its information, the emotion engine analyzes the context and emotions of the input and automatically completes the information.
[1372] This system enables efficient measurement of the number of species and population size of organisms, improving analysis accuracy and optimizing the user experience.
[1373] The following describes the processing flow.
[1374] Step 1: Data Collection
[1375] The server periodically collects video footage from multiple cameras installed at the observation site.
[1376] The server stores the collected video data in a temporary storage folder and simultaneously saves it as a backup in the central database.
[1377] For example, video footage is collected from observation points A and B every hour and stored in a central database.
[1378] Step 2: Video Analysis
[1379] The server reads unanalyzed video data from the central database.
[1380] The server loads a pre-trained deep learning model (e.g., YOLO or Mask R-CNN) and analyzes the video data.
[1381] The server assigns bounding boxes, species labels, and confidence scores to organisms detected in the video. For example, it identifies three deer and two rabbits from video footage of observation point A.
[1382] Step 3: Individual Identification
[1383] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species.
[1384] The server assigns a unique ID to each individual and identifies them by considering their distinctive physical characteristics (patterns, body size, etc.). For example, if one of three deer has a different pattern, the server will use that as the basis for individual identification.
[1385] Step 4: Database Update
[1386] The server stores the species and individual identification information of the identified organisms in a database.
[1387] The server adds data of unidentified organisms to an unidentified list and saves the video footage for later use as training data. For example, it adds newly discovered insects to the unidentified list.
[1388] Step 5: Register training data and retrain the AI model
[1389] The user checks the unidentified list and enters the species name and information of the newly identified organism.
[1390] The server updates the training data based on new biological information provided by the user and prepares to retrain the AI model.
[1391] The server retrains to improve the recognition accuracy of the AI model.
[1392] Step 6: Density Calculation
[1393] The server compiles the number of organisms in each observation area based on the analysis results.
[1394] The server uses the aggregated number of organisms to determine the biodensity (number of individuals / m²) for each area. 2 Calculate the density of deer at observation point A. For example, the density of deer at observation point A is 1.5 deer / m². 2 Rabbit density is 1 rabbit / m 2 This is the result.
[1395] Step 7: Report Generation
[1396] The server generates reports for each observation area based on the calculated biodensity.
[1397] The server saves the generated reports to a central database.
[1398] Step 8: Displaying the User Interface
[1399] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user.
[1400] Users review observation results through the interface and input information about organisms that were not identified. This information is then used to train the AI model for the next time.
[1401] Step 9: Combining Emotional Engines
[1402] The server uses an emotion engine to recognize the user's emotions. For example, it analyzes sensor data from webcams and microphones to evaluate the user's facial expressions and tone of voice.
[1403] The server customizes how analysis results and reports are displayed based on emotional data. For example, if a user is experiencing stress, it adjusts the frequency and level of detail of information provided to reduce the user's burden.
[1404] Users also use the sentiment engine to analyze the context and emotions of the unidentified organisms they input. This allows for simple corrections and automatic completion, ensuring that accurate information is reflected in the database.
[1405] Through the steps described above, this system can efficiently measure the number of species and population size of organisms, improve analysis accuracy year after year, and optimize the user experience.
[1406] (Example 2)
[1407] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1408] Conventional biological observation systems can identify species when analyzing video footage from observation sites, but they lack the means to uniquely identify individual organisms. Furthermore, they fail to provide information that takes into account the user's emotional state, resulting in an unoptimized user experience. Therefore, improvements in both observation accuracy and user experience are needed.
[1409] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1410] In this invention, the server includes means for acquiring video data from multiple cameras installed at observation sites, means for analyzing the acquired video data to detect and identify organisms in the video, means for storing data of detected and identified organisms in a storage device, means for registering data of organisms that were not identified as training data and retraining an artificial intelligence model, means for calculating the organism density for each observation site and generating a report, means for uniquely identifying each individual organism using the artificial intelligence model, and means for recognizing the user's emotions and customizing the analysis results and reports based on that data. This improves the accuracy of organism observation and optimizes the user experience.
[1411] A "camera" is a device installed at an observation point to acquire video data.
[1412] "Video data" refers to video information acquired from multiple recording devices.
[1413] A "server" is a computing device that analyzes video data to detect and identify living organisms.
[1414] A "memory device" is a data storage device used to store data on detected and identified organisms.
[1415] An "artificial intelligence model" is a general term for machine learning algorithms used for detecting and identifying living organisms, as well as for uniquely identifying individuals.
[1416] "Retraining" is a learning process that uses data from organisms that were not identified to improve the accuracy of an artificial intelligence model.
[1417] "Biodensity" refers to the number of organisms per unit area at a specific observation point.
[1418] A "report" is a document that summarizes analysis results, biodensities, and other data.
[1419] "User emotions" refer to the psychological state a user experiences while using a system, such as interest or stress.
[1420] "Customization" refers to adjusting how analysis results and reports are displayed based on user sentiment data.
[1421] This invention is a system that acquires video data from multiple cameras installed at a specific observation point, analyzes that data to identify species and individuals, and calculates biodensity. Furthermore, it is characterized by its ability to improve the user experience by incorporating an emotion engine that recognizes the user's emotions. This system has a configuration centered on a server, terminals, and users, and operates as follows.
[1422] Data collection
[1423] The server periodically collects video data from multiple cameras installed at observation sites and stores it in a central database. For example, it collects video data from observation sites A and B every hour and stores it in the central database. It downloads and manages video streams using HTTP requests.
[1424] Video analysis
[1425] The server loads a pre-trained deep learning model (such as YOLO or Mask R-CNN) to analyze the acquired video data. This model is used to analyze each frame of the video, and an object detection algorithm is applied. The GPU is used to perform model inference at high speed, detecting the location and species of organisms in the video. For example, analyzing video from observation point A can detect three deer and two rabbits in each frame.
[1426] Individual identification
[1427] The server applies additional deep learning models and clustering algorithms to each detected organism to identify individuals within the same species. Each individual is assigned a unique ID and identified based on characteristics such as markings and body size. For example, if one of three deer has a different marking, it is identified as "Deer 1," "Deer 2," and "Deer 3" based on that.
[1428] Database update
[1429] The server updates the database with species and individual identification information for identified organisms. It uses SQL queries to insert the identification information into the appropriate tables. Data for unidentified organisms is added to an unidentified list, and their videos are saved separately. For example, a newly discovered insect is added to the unidentified list, and its video is saved to cloud storage.
[1430] Registration of training data and retraining of the AI model
[1431] The user reviews the unidentified list and enters the species name and other information of newly identified organisms. This operation can be performed through a form in the web application. This information is sent to the server in JSON format. The server incorporates the new organism data provided by the user into the training database and generates a new training set. The AI model is retrained using a GPU cluster to improve accuracy.
[1432] density calculation
[1433] The server aggregates the number of organisms in each observation area based on the analysis results and calculates the biodensity. It obtains area information for each observation point from the database and calculates the density by dividing the number of individuals by the area. For example, using data obtained from the scan of observation point A, it calculates that "the density of deer is 1.5 individuals / m²". 2 "The rabbit density is 1 rabbit / m²" 2 The calculation is as follows:
[1434] Report generation
[1435] The server generates reports for each observation area based on the calculated biodensity. A Python script is used to automatically generate these reports, which are then stored in a central database. These reports include density information and individual identification information for each species.
[1436] User Interface Display
[1437] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user. This interface is designed using front-end technologies such as JavaScript and React. Users can view results for specific observation points and input information about organisms that were not identified. This input information is sent to the server and used for training the next AI model.
[1438] Combination of emotional engines
[1439] The server uses an emotion engine to recognize the user's emotions. It analyzes user operation logs and input data and applies algorithms to predict the emotional state. Emotional data is reflected in report displays and alert generation. For example, if the user is stressed, the frequency of information provided may be reduced or the display mode may be switched to a simplified mode. Even when the user enters information about an unfamiliar insect, the emotion engine analyzes the context and emotions and performs automatic completion or correction.
[1440] Specific example
[1441] For example, based on video data acquired from observation point A, the server identifies three deer and two rabbits and assigns a unique ID to each individual. Subsequently, when the user uses a device to view the analysis results, the emotion engine recognizes the user's interests and stress levels and customizes the displayed content. When the user discovers a new unidentified insect and enters its information, the emotion engine analyzes the context and emotions of the input and automatically completes the information.
[1442] Example of a prompt
[1443] Analyze the following video footage and identify the species of organisms shown. Also, identify each individual organism and update the database with your findings.
[1444] For organisms on the unidentified list, please retrain the AI model based on the information entered by the user to improve the identification accuracy for the next time.
[1445] Analyze user sentiment and customize how the analysis results and reports are displayed accordingly.
[1446] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1447] Step 1:
[1448] The server periodically collects video data from multiple cameras installed at observation sites. Its input is the video stream from the observation sites, obtained via HTTP requests. Specifically, it uses a scheduling function to access the cameras every hour and download the video data. The output is the storage of the acquired video data in a central database.
[1449] Step 2:
[1450] The server loads a pre-trained deep learning model (e.g., YOLO or Mask R-CNN) to analyze the acquired video data. The input is the stored video data. Specifically, each frame of the video is divided, and an object detection algorithm is applied. The output is the location information and species identification result of organisms in each frame. Model inference is performed at high speed using a GPU, and analysis is performed efficiently.
[1451] Step 3:
[1452] The server applies additional deep learning models and clustering algorithms to the detected organisms to identify individuals within the same species. The input is the location information and species identification results of the organisms obtained in step 2. Specifically, a unique ID is assigned to each individual, and features such as patterns and body size are analyzed. The output is the uniquely identified individual information. This makes it possible to make specific identifications such as "Deer 1," "Deer 2," and "Deer 3."
[1453] Step 4:
[1454] The server updates the database with the species and individual identification information of the identified organisms. The input is the unique individual information obtained in step 3. Specifically, it inserts this information into the appropriate table in the database using an SQL query. The output is the updated database information. Data of organisms that were not identified is added to the unidentified list, and their video data is saved to cloud storage.
[1455] Step 5:
[1456] The user reviews the unidentified list and enters the species name and other information of newly identified organisms. This input is provided through a web application form. This information is sent to the server in JSON format. The server incorporates this new data into the training database and generates a new training set. The output is an updated training data set and a new training set to be used for the next training session.
[1457] Step 6:
[1458] The server aggregates the number of organisms in each observation area based on the analysis results and calculates the biodensity. The input is the database information updated in step 4. Specifically, it obtains area information for each observation point from the database and calculates the density by dividing the number of individuals by the area. The output is the biodensity information for each observation area. For example, from the data for observation point A, it might say "the deer density is 1.5 individuals / m²". 2 "The rabbit density is 1 rabbit / m²" 2 The calculation is as follows:
[1459] Step 7:
[1460] The server generates reports for each observation area based on biodensity. The input is the biodensity information obtained in step 6. A Python script is used to automatically generate the reports and save them to the central database. The output is the generated report. This report includes density information and individual identification information for each species.
[1461] Step 8:
[1462] The terminal provides an interface that displays analysis results and biodensity reports when accessed by the user. Input consists of reports and analysis results obtained from the server. The web page is designed using front-end technologies such as JavaScript and React. Output consists of analysis results and reports displayed to the user. The user can review results for specific observation areas and add information about organisms that were not identified.
[1463] Step 9:
[1464] The server uses an emotion engine to recognize the user's emotions. Input is the user's operation logs and input data. An algorithm that predicts emotional states is used to detect stress, interest, etc. The output is the recognized emotion information. Based on this, the frequency of information provision and the display method are automatically customized. Even when the user inputs information about an unknown insect, the emotion engine analyzes the context and emotions and automatically completes or corrects the information.
[1465] (Application Example 2)
[1466] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1467] There is a need to accurately understand customer behavior and emotions in stores and use that information to improve product placement and service quality, but current technology makes it difficult to fully achieve this. Furthermore, there is a lack of efficient methods for calculating and identifying biological densities. Therefore, more advanced data analysis technologies are needed to improve the efficiency of store operations and enhance the customer experience.
[1468] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1469] In this invention, the server includes means for acquiring video footage from multiple cameras installed at observation points, means for analyzing the acquired video footage to detect and identify organisms in the footage, means for storing data of the detected and identified organisms in a database, means for registering data of organisms that were not identified as training data and retraining an AI model, means for calculating the organism density for each observation point and generating a report, means for filming and analyzing customer behavior from multiple cameras installed in the store, means for analyzing the video footage to identify customers and assign a unique ID, and means for identifying emotional data using an emotion engine that analyzes customer emotions. This makes it possible to analyze customer behavior and emotions in detail, optimize store operations based on this analysis, and improve the customer experience.
[1470] An "observation point" is a specific area or location where the system acquires video footage and uses it for analysis.
[1471] A "camera" is a device used to capture video footage and collect that data.
[1472] A "server" is a central processing unit used to analyze collected video footage, store data, and retrain AI models.
[1473] "Video footage" is a media format composed of a series of still images that records the situation at an observation point or inside a store in real time.
[1474] "Living organisms" refers to species and individuals of animals, plants, insects, and other organisms present at the observation site.
[1475] "Identification" refers to the act of identifying a specific species, individual, or customer from analyzed video footage.
[1476] A "database" is an information system for systematically storing collected and analyzed data.
[1477] An "AI model" is an algorithm that uses deep learning or machine learning to perform data analysis.
[1478] "Retraining" is the process of improving the accuracy of an AI model using new training data.
[1479] "Biological density" is an index calculated by determining the number of organisms in a specific observation location per unit area.
[1480] A "report" is a quantitative and qualitative report generated based on the results of an analysis.
[1481] A "customer" is a person who engages in activities within a store.
[1482] An "emotion engine" is an analytical device and software that estimates emotions from a customer's facial expressions and body movements.
[1483] A "unique ID" is identification information that assigns a unique identifier to each identified organism or customer.
[1484] This invention is a system that acquires video footage from multiple cameras installed at specific observation points, analyzes the data to identify biological species and individuals, and calculates biodensity. It also aims to improve the user experience by incorporating an emotion engine that recognizes user emotions. Furthermore, by applying this system to physical stores and analyzing customer behavior and emotions, it aims to optimize store operations and improve the quality of customer service.
[1485] 1. Data Collection
[1486] The server periodically collects video footage from multiple cameras installed at observation points and stores it in a central database. For example, it collects video from observation points A and B every hour and stores it in the central database. It also periodically acquires video data from cameras installed inside the store. This allows for continuous monitoring of customer behavior.
[1487] 2. Video Analysis
[1488] The server uses the following AI model to analyze the acquired video footage.
[1489] Object detection algorithms such as YOLO (You Only Look Once) or Mask R-CNN
[1490] These models are used to detect organisms in video footage and identify their species and individual characteristics. Additionally, in in-store video footage, customers are identified, and each customer is assigned a unique ID.
[1491] 3. Individual Identification
[1492] The server uses additional deep learning models and clustering algorithms to uniquely identify each individual within the same species or each customer. For example, it can identify individuals based on unique physical characteristics or body movements.
[1493] 4. Database Update
[1494] The server stores the species and individual identification information of identified organisms in a database. Unidentified data is added to an unidentified list, and its video is saved for later use as training data. The same applies to customers, storing customer behavior data and sentiment data in the database.
[1495] 5. Registering training data and retraining the AI model
[1496] Users can review the unidentified list and input the species name and information of newly identified organisms. Based on this, the server retrains the AI model. Similarly, for customer data, the model is retrained based on customer behavior patterns and sentiment data.
[1497] 6. Emotion analysis
[1498] The server uses an emotion engine to recognize the user's emotions. Based on this emotion data, it customizes how analysis results and reports are displayed. For example, if a user is feeling stressed in a store, the frequency and level of detail of information provided will be adjusted. It is also possible to analyze customer emotion data and provide appropriate services.
[1499] 7. Report generation and user interface display
[1500] The server generates reports for each observation area and store based on calculations of biodensity and customer behavior. These reports are stored in a central database and accessible to users via terminals. Users can review the analysis results and reports and input information about unidentified organisms and customers.
[1501] Specific example
[1502] For example, based on video footage acquired from observation point A, the server identifies three deer and two rabbits and assigns a unique ID to each individual. When the user then reviews the analysis results, the emotion engine recognizes the user's interests and stress levels and customizes the displayed content. Furthermore, when the user discovers a new unidentified insect and inputs its information, the emotion engine analyzes the context and emotions of the input and automatically completes the information.
[1503] Example of a prompt
[1504] "Please analyze the behavior of customers who spend extended periods of time in specific areas of the store (e.g., the cosmetics section), and based on the results of this analysis, including emotional data, please propose improvements."
[1505] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1506] Step 1:
[1507] Data collection
[1508] The server periodically collects video footage from multiple cameras installed at observation points and within the store. For example, the server collects video from observation points A and B every hour and stores it in a central database. It also acquires video data from cameras inside the store and stores it in the central database as well.
[1509] Input: Camera footage
[1510] Output: Video data stored in the central database
[1511] Step 2:
[1512] Video analysis
[1513] The server uses an AI model to analyze the acquired video footage. Specifically, it employs object detection algorithms such as YOLO and Mask R-CNN. This allows it to detect living beings and customers in the video and assign a unique ID to each individual or customer.
[1514] Input: Video data
[1515] Output: Identified individual / customer data (with unique ID)
[1516] Step 3:
[1517] Individual identification
[1518] The server further analyzes the detailed characteristics of identified organisms and customers using deep learning models and clustering algorithms. It uniquely identifies each individual or customer based on their physical characteristics and behavioral patterns.
[1519] Input: Identified individual / customer data
[1520] Output: Detailed data on individual customers (feature-based unique IDs)
[1521] Step 4:
[1522] Database update
[1523] The server stores data on identified organisms and customers in a database. Data that is not identified is added to an unidentified list, and this video is saved and used as training data later.
[1524] Input: Individual / customer detailed data, unidentified data
[1525] Output: Updated database
[1526] Step 5:
[1527] Registration of training data and retraining of the AI model
[1528] The user reviews the unidentified list and inputs information about newly identified organisms or customers. Based on this information, the server retrains the AI model. New training data is added during this process, improving the model's accuracy.
[1529] Input: Unidentified list, user input data
[1530] Output: Retrained AI model
[1531] Step 6:
[1532] Emotion analysis
[1533] The server uses an emotion engine to recognize the emotions of customers and users. For example, it estimates emotions from a customer's facial expressions and body movements and stores them as emotion data. Based on this, it customizes how analysis results and reports are displayed.
[1534] Input: Video data
[1535] Output: Sentiment data
[1536] Step 7:
[1537] Report generation and user interface display
[1538] The server generates reports for each observation area and store based on calculations of biodensity, customer behavior, and emotions. These reports are stored in a central database and can be accessed by users via terminals. Users can review the analysis results and reports and input information about unidentified organisms and customers.
[1539] Input: Calculation results, biological / customer data, sentiment data
[1540] Output: Biodensity report, customer behavior report, user interface display
[1541] Example of a prompt
[1542] "Please analyze the behavior of customers who spend extended periods of time in specific areas of the store (e.g., the cosmetics section), and based on the results of this analysis, including emotional data, please propose improvements."
[1543] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1544] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1545] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1546] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1547] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1548] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1549] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1550] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1551] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1552] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1553] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1554] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1555] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1556] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1557] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1558] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1559] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1560] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1561] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1562] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1563] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1564] The following is further disclosed regarding the embodiments described above.
[1565] (Claim 1)
[1566] A means of acquiring video footage from multiple cameras installed at the observation site,
[1567] A means for analyzing acquired video footage to detect and identify organisms in the footage,
[1568] Means for storing data of detected and identified organisms in a database,
[1569] A method for registering data on unidentified organisms as training data and retraining the AI model,
[1570] A system that includes means for calculating the biodensity of each observation point and generating a report.
[1571] (Claim 2)
[1572] The system according to claim 1, comprising means for providing a generated report to a user through a user interface and for the user to input information on unidentified organisms.
[1573] (Claim 3)
[1574] The system according to claim 1, further comprising means for uniquely identifying each individual organism in addition to species identification of organisms detected in the video.
[1575] "Example 1"
[1576] (Claim 1)
[1577] A means for acquiring video data from multiple image acquisition devices installed at the observation site,
[1578] Means for storing acquired video data in a data storage device,
[1579] A means for inputting stored video data into an analysis device for analysis and automatically detecting and identifying organisms in the video,
[1580] A means for storing data of identified organisms in a database,
[1581] A means of registering data of organisms that were not identified as training data and retraining the recognition model device,
[1582] A system that includes means for calculating biodensity at each observation point and generating report data.
[1583] (Claim 2)
[1584] The system according to claim 1, comprising means for providing generated report data to a user through a user interface and enabling the user to input information on unidentified organisms.
[1585] (Claim 3)
[1586] The system according to claim 1, further comprising means for uniquely identifying each individual organism in addition to species identification of organisms detected in the video.
[1587] "Application Example 1"
[1588] (Claim 1)
[1589] A means of acquiring video footage from multiple cameras installed at the observation site,
[1590] A means for analyzing acquired video footage to detect and identify animals in the footage,
[1591] Means for storing data of detected and identified animals in a storage device,
[1592] A method for registering data on animals that were not identified as training data and retraining the AI model,
[1593] A means for calculating the animal density at each observation point and generating a report,
[1594] A means of analyzing acquired video footage to identify specific individuals and their density in real time,
[1595] Means for storing data of detected and identified persons in a database,
[1596] A system that includes a means for calculating the density of people in each area based on the analysis results.
[1597] (Claim 2)
[1598] The system according to claim 1, comprising means for providing a generated report to a user through a user interface, means for the user to input information on an unidentified animal, and means for issuing an alert based on an identified person.
[1599] (Claim 3)
[1600] The system according to claim 1, further comprising means for uniquely identifying each individual animal, in addition to identifying the species of animals detected in the video, and means for uniquely identifying a specific person.
[1601] "Example 2 of combining an emotion engine"
[1602] (Claim 1)
[1603] A means of acquiring video data from multiple cameras installed at the observation site,
[1604] A means for analyzing acquired video data to detect and identify organisms in the video,
[1605] Means for storing data of detected and identified organisms in a storage device,
[1606] A method for registering data of unidentified organisms as training data and retraining an artificial intelligence model,
[1607] A means for calculating the biodensity of each observation point and generating a report,
[1608] A means of uniquely identifying each individual organism using an artificial...
Claims
1. A means of acquiring video footage from multiple cameras installed at the observation site, A means for analyzing acquired video footage to detect and identify organisms in the footage, Means for storing data of detected and identified organisms in a database, A method for registering data on unidentified organisms as training data and retraining the AI model, A system that includes means for calculating the biodensity of each observation point and generating a report.
2. The system according to claim 1, further comprising means for providing the generated report to the user through a user interface and for the user to input information on unidentified organisms.
3. The system according to claim 1, further comprising means for uniquely identifying each individual organism in addition to species identification of organisms detected in the video.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A