System

The system addresses the challenge of identifying and responding to vermin and insects by using cameras and microphones for data collection, analysis, and automated countermeasures, ensuring swift and effective safety measures.

JP2026017992APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119053
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Conventional methods struggle to quickly and effectively identify and respond to the presence of vermin and insects in urban areas, posing safety risks to residents, due to difficulties in target identification and the need for manual operation.

Method used

A system that uses cameras and microphones to collect video and audio data, analyzes it to identify targets, automatically generates appropriate countermeasures, and executes them, including features like loud sounds or ultrasonic waves, with continuous monitoring and reanalysis for effectiveness.

Benefits of technology

Enables early detection and rapid, tailored responses to vermin and insects, ensuring a safe living environment without requiring special technical knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017992000001_ABST
    Figure 2026017992000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for collecting video and audio using a camera and a microphone installed in an area where pests appear; means for receiving the collected video and audio data; means for analyzing the received data and identifying an object; means for automatically generating a coping method corresponding to the identified object; means for transmitting a command for executing the generated coping method; and means for executing the coping method based on the received command.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The "problem that the invention aims to solve" and "means for solving the problem" are described below.

[0005] patent specification

[0006] In recent years, even in urban areas, there has been an increase in the number of cases where vermin and insects such as bears and hornets have appeared, putting residents at risk of injury and even death. This has resulted in problems that threaten safe living environments. To solve this problem, it is necessary to detect the presence of vermin and insects early before they appear and implement appropriate countermeasures. However, conventional methods often have difficulty identifying the target and selecting appropriate countermeasures, making it difficult to respond quickly and effectively. The present invention aims to provide a system that prevents the appearance of vermin and insects and provides a safe living environment. [Means for solving the problem]

[0007] The system of the present invention has the following features: It includes a means for collecting video and audio using cameras and microphones installed in areas where pests or vermin appear, a means for receiving the collected video and audio data, a means for analyzing the received data and identifying the target, a means for automatically generating a countermeasure for the identified target, a means for sending a command to execute the generated countermeasure, a means for executing the countermeasure based on the received command, and a means for re-monitoring and analyzing the situation after the countermeasure is executed. This system enables early detection of targets and appropriate countermeasures, thereby ensuring the safety of residents. Furthermore, by automatically generating a countermeasure using appropriate sound, light, or video depending on the target, it is possible to provide a system that can respond to a variety of situations.

[0008] ---

[0009] A "vermin" is an animal that appears in urban or residential areas and has the potential to cause harm to residents.

[0010] "Pests" are insects that infest urban and residential areas and can cause harm to people.

[0011] A "camera" is a device for capturing images and recording or transmitting the data.

[0012] A "microphone" is a device for collecting sound and recording or transmitting that data.

[0013] "Video" is visual data captured by a device such as a camera.

[0014] "Audio" is sound data collected by a device such as a microphone.

[0015] "Data" means electromagnetic or electronic records containing visual or audio information.

[0016] "Receiving" is the operation or process of receiving transmitted data.

[0017] "Analysis" is the process of examining received data in detail to understand its content and characteristics.

[0018] A "target" is a specific entity detected as a result of analysis, such as a vermin or insect.

[0019] "Identification" is the process of identifying an object individually through analysis and clarifying its type and characteristics.

[0020] "Countermeasures" are specific measures or methods taken to avoid the existence of an identified object.

[0021] "Automatic generation" is the process by which a system or device autonomously creates methods or means without manual operation.

[0022] A "command" is a specific instruction or command to carry out a countermeasure.

[0023] "Transmitting" is the operation or process of sending data or commands to another device or system.

[0024] "Execution" is the process of specifically taking action based on the command sent.

[0025] "Monitoring" is the process of continuously observing a particular area or situation to detect changes or anomalies.

[0026] "Situation" refers to the current state or environment of the monitored area or object.

[0027] "Repeated" refers to the repetition of a process or operation that has been performed once, and indicates a process that includes this. [Brief explanation of the drawings]

[0028] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0029] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0030] First, the terms used in the following description will be explained.

[0031] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0032] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0033] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0034] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0035] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0036] [First embodiment]

[0037] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0038] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0039] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0040] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0041] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0042] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0043] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0044] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0045] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0046] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0047] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0048] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0049] ---

[0050] MODE FOR CARRYING OUT THE INVENTION

[0051] The present invention relates to a system for detecting the appearance of pests and vermin at an early stage and automatically taking appropriate measures. This system includes a series of processes that use a camera and a microphone to collect video and audio within an area, analyze the video and audio to identify the target, and take appropriate measures. The following describes in detail an embodiment of the present invention.

[0052] Overall system overview

[0053] This system consists of cameras and microphones (hereafter referred to as terminals) installed within the area, and a server that receives and analyzes the data sent from these terminals and implements countermeasures.

[0054] Program processing

[0055] 1. Information gathering

[0056] The device captures video and audio within the set area in real time, and this data is used as the basis for understanding the situation within the area.

[0057] 2. Data Transmission

[0058] The device transmits the captured video and audio data to a server via a network, enabling real-time data collection and transmission.

[0059] 3. Data Analysis

[0060] The server analyzes the received video and audio data, and performs preprocessing such as noise removal and data normalization.

[0061] The server uses an AI model (e.g., an image recognition model or a voice recognition model using deep learning) to identify the object (e.g., a bear or a hornet).

[0062] 4. Generating solutions

[0063] The server automatically generates the most effective response to the object identified through the analysis: for example, a loud cracking sound for a bear, or ultrasonic waves for a hornet.

[0064] 5. Sending solutions

[0065] The server sends the generated command for the solution to the terminal, which includes the specific solution and the timing for its execution.

[0066] 6. Implementing the solution

[0067] The device will then take action based on the command it receives, for example, playing a specific sound or emitting a bright light.

[0068] 7. Recheck and Readdress

[0069] After the countermeasure is implemented, the device continues to monitor the situation within the area and captures video and audio again, which are then sent to the server for further analysis.

[0070] If the server determines that further action is necessary, it generates an appropriate countermeasure and sends it to the device. By repeating this process, the target object is completely avoided.

[0071] Specific examples

[0072] For example, if a bear appears near a house, the system operates as follows:

[0073] The terminals (cameras installed around the home) capture images of bears and send the data to a server.

[0074] The server preprocesses the received footage and uses an AI model to identify the bear.

[0075] The server generates a command to play a loud cracking sound that is effective against bears and sends it to the device.

[0076] The device plays a loud cracking sound to scare off bears.

[0077] The device will capture the video again to check if the bear has completely left the area, and take further action if necessary.

[0078] Similarly, if a hornet's nest is discovered on the eaves of a house, the device captures video and audio, and the server analyzes them to identify the nest. The server selects a countermeasure to generate ultrasonic waves that repel hornets and sends it to the device. The device repels hornets by emitting ultrasonic waves.

[0079] This system allows users to detect the appearance of pests and vermin early and deal with them effectively, thereby providing a safe living environment. This series of processes is fully automated, so users do not need any special technical knowledge or operation to operate the system.

[0080] The above is a specific embodiment for carrying out the present invention.

[0081] The processing flow will be explained below.

[0082] Step-by-step process

[0083] Step 1:

[0084] The device captures video and audio within the set area in real time. Video is captured using a camera and audio is collected using a microphone. The collected data is temporarily stored in local storage and updated periodically.

[0085] Step 2:

[0086] The devices transmit the collected video and audio data over a network to a server. This data is streamed in real time or sent in batches. The data transmission is encrypted to ensure security.

[0087] Step 3:

[0088] The server receives the video and audio data sent from the device, saves the data in the data storage, and copies it to the workspace for analysis.

[0089] Step 4:

[0090] The server performs preprocessing on the received data, specifically noise removal, data normalization, and missing value imputation, to improve data quality and enhance analysis accuracy.

[0091] Step 5:

[0092] The server inputs the preprocessed data into the AI ​​model to detect objects. It uses an image recognition model to identify objects from video data and a voice recognition model to detect abnormal sounds from audio data.

[0093] Step 6:

[0094] The server automatically generates the most appropriate response to the object identified through the analysis: for example, a loud cracking sound is generated if a bear is identified, and ultrasonic waves are generated if a hornet is identified.

[0095] Step 7:

[0096] The server sends the generated command for the countermeasure to the terminal. This command includes the countermeasure to be executed and its parameters (e.g., type and intensity of sound, light pattern, etc.).

[0097] Step 8:

[0098] The device will then take action based on the command it receives, for example, playing a repellent sound from a specific speaker if a voice command is received, or emitting a bright light if a video command is received.

[0099] Step 9:

[0100] The device will monitor the situation again after the countermeasure is implemented, capturing video and audio data, which will also be sent to the server again.

[0101] Step 10:

[0102] The server analyzes the retransmitted data and determines whether additional action is required. If the object is still present, it generates an appropriate action again and sends it to the device. Repeat from step 6 if necessary.

[0103] The above is the flow of the program processing including the specific operations at each step.

[0104] Example 1

[0105] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0106] Conventional systems for detecting and responding to the emergence of pests and vermin have problems with real-time accuracy, often resulting in delayed responses and insufficient effectiveness. Furthermore, manual operation is required, and users without special technical knowledge have had the problem of being unable to operate the systems properly. Furthermore, the methods offered as avoidance measures by the systems are uniform, making it difficult to provide optimal countermeasures tailored to the characteristics of the target.

[0107] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0108] In this invention, the server includes a means for analyzing collected video and audio data to identify an object, a means for automatically generating a countermeasure corresponding to the identified object, and a means for transmitting a command for executing the generated countermeasure, thereby enabling high-precision detection of an object in real time and automatic appropriate countermeasures.

[0109] An "image capture device" is a device installed to capture video within a specific area in real time.

[0110] "Audio recording device" refers to equipment installed to record audio within a specific area.

[0111] "Means for collecting" refers to the function or process of acquiring video and audio data using image capture devices and audio recording devices.

[0112] The "means for receiving" is a function or process for transmitting collected video and audio data to a server via a network and for the server to receive it.

[0113] "Means for analyzing" refers to algorithms or systems that process received video and audio data and identify objects.

[0114] The "means for automatic generation" is a function or process that automatically creates an appropriate countermeasure for the target object based on the results of data analysis.

[0115] The "means for sending a command" is a function or process that sends an instruction to execute the generated countermeasure to the terminal via the network.

[0116] The "means for executing" refers to a function or device that the terminal actually physically executes the countermeasure based on the received command.

[0117] "Means for monitoring and analysis" refers to the function or process of rechecking the situation after the countermeasures have been implemented and reanalyzing it based on newly collected data.

[0118] The "pre-processing means" refers to a function or process that performs noise removal and data normalization on the collected video and audio data.

[0119] A "generative AI model" is a model or algorithm that uses deep learning or machine learning to identify objects from data.

[0120] MODE FOR CARRYING OUT THE INVENTION

[0121] The present invention provides a system for detecting the appearance of pests and vermin at an early stage and automatically taking appropriate countermeasures. This system includes a series of processes that use an image capture device and an audio recording device to collect video and audio within an area, analyze the video and audio, identify the target, and take appropriate countermeasures. The following describes in detail the embodiments of the present invention.

[0122] Overall system overview

[0123] This system consists of image capture devices and audio recording devices (hereafter referred to as terminals) installed within the area, and a server that receives and analyzes the data sent from these terminals and executes countermeasures. The system is designed to be easy to operate, even if the user does not have special technical knowledge.

[0124] Program processing

[0125] 1. Information collection via devices

[0126] The device captures video and audio within the set area in real time. The image capture device captures video at 30 frames per second, and the audio recorder records audio at 24-bit / 48kHz.

[0127] 2. Data transmission by the terminal

[0128] The device periodically compresses the captured video and audio data, converts it into a data structure such as JSON or Protobuf format, and sends it to the server using HTTP / HTTPS or WebSocket protocols.

[0129] 3. Data analysis by the server (preprocessing)

[0130] The server first decodes the received data and performs noise reduction on the video and audio. The video data is decoded using the OpenCV library to reduce blur and noise, and the audio data is decoded using Python's pydub and SciPy.

[0131] 4. Data analysis by the server (identification of the target object)

[0132] The server analyzes the video and audio data using a generative AI model (for example, a deep learning model using TensorFlow or PyTorch) and uses object detection algorithms such as YOLO (You Only Look Once) or R-CNN to identify objects such as bears or hornets.

[0133] 5. Server-generated solutions

[0134] Based on the identified object, the server uses a pre-defined rules-based system to generate the optimal response, such as playing a loud cracking sound for a bear or generating ultrasound for a hornet.

[0135] 6. Server sends solution

[0136] The server sends the generated command for the solution to the appropriate device at the appropriate time. This command includes the specific content of the solution (e.g., the path of the audio file, playback timing, etc.).

[0137] 7. Implementing measures on the device

[0138] The device will then take action based on the command it receives, such as playing a loud crackling sound through the speaker or flashing an LED light.

[0139] 8. Reconfirmation and re-action by terminal

[0140] After implementing a countermeasure, the device continues to monitor the situation within the area and recaptures video and audio data. This data is then sent to the server for further analysis. If the server determines that further countermeasures are necessary, it generates a new appropriate countermeasure and sends it to the device. This process is repeated as many times as necessary.

[0141] Specific examples

[0142] For example, if a bear appears near a house, the system operates as follows:

[0143] 1. Information collection via devices

[0144] The device captures video at 30 frames per second using an image capture device installed in the garden, and records audio at 24-bit / 48kHz using an audio recording device.

[0145] 2. Data transmission by the terminal

[0146] The device compresses the captured data every 5 seconds, converts it into JSON format, and sends it to the server using the HTTP protocol.

[0147] 3. Data analysis by the server (preprocessing)

[0148] The server uses OpenCV to remove noise from the received video data and pydub to perform noise filtering on the audio data.

[0149] 4. Data analysis by the server (identification of the target object)

[0150] The server uses the YOLO model to analyze the video data and identify the presence of bears.

[0151] 5. Server-generated solutions

[0152] The server generates commands based on pre-defined rules to play a loud cracking sound that works on bears.

[0153] 6. Server sends solution

[0154] The server sends the generated command (e.g., the path to the audio file and the playback timing) to the terminal.

[0155] 7. Implementing measures on the device

[0156] Based on the command received, the device plays a loud cracking sound through its speaker to scare off bears.

[0157] 8. Reconfirmation and re-action by terminal

[0158] The device captures the video again to check if the bear has completely left the area. If the bear is still in the area, the server generates a new countermeasure and sends it to the device to execute again.

[0159] Example prompts for generative AI models

[0160] "Please tell me how to use this system to automatically detect the appearance of bears and take appropriate action."

[0161] This is a specific mode for carrying out the invention.

[0162] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0163] Program processing steps

[0164] Step 1: Gather information

[0165] The device captures video and audio within the area in real time. Specifically, the image capture device captures video at 30 frames per second, and the audio recording device records audio at 24-bit / 48kHz.

[0166] Input: Physical conditions within the set area

[0167] Output: Video and audio data

[0168] Step 2: Send data

[0169] The device periodically compresses the captured video and audio data, converts it into JSON or Protobuf format, and sends it to the server using HTTP / HTTPS or WebSocket protocol. For example, the device compresses and sends the data every 5 seconds.

[0170] Input: Video and audio data

[0171] Output: Compressed data

[0172] Step 3: Data analysis (preprocessing)

[0173] The server decodes the received data and performs noise reduction on the video and audio. Specifically, the OpenCV library is used to reduce blur and noise in the video data, and pydub and SciPy are used to perform noise filtering on the audio data.

[0174] Input: Compressed data

[0175] Output: Denoised data

[0176] Step 4: Data analysis (target identification)

[0177] The server analyzes the video and audio data using a deep learning model (e.g., a generative AI model using TensorFlow or PyTorch) to identify objects using object detection algorithms such as YOLO or R-CNN.

[0178] Input: Denoised data

[0179] Output: Information about the identified object (e.g., location and type of bear or hornet)

[0180] Step 5: Generate solutions

[0181] Based on the identified object, the server uses a rules engine to automatically generate the optimal response, for example, a loud cracking sound for a bear or ultrasonic waves for a hornet.

[0182] Input: Information about the identified object

[0183] Output: Action command

[0184] Step 6: Submit your solution

[0185] The server generates a command for the appropriate action and sends it to the device at the appropriate time, which includes specific instructions such as the path to the audio file and the timing of playback.

[0186] Input: Solution command

[0187] Output: Commands sent to the terminal

[0188] Step 7: Implementing the solution

[0189] The device will then take action based on the command it receives, such as playing a loud crackling sound through the speaker or flashing an LED light.

[0190] Input: Command sent to the terminal

[0191] Output: Actions taken

[0192] Step 8: Reassess and readdress

[0193] Even after the device has implemented the countermeasures, it continues to monitor the situation within the area, recapturing video and audio data and sending it to the server. If it determines that reanalysis is necessary, it generates an appropriate countermeasure again and sends it to the device.

[0194] Input: Post-execution video and audio data

[0195] Output: Updated workaround command

[0196] The above is the processing flow of the specific program of the system.

[0197] (Application example 1)

[0198] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0199] Automating the early detection of pests and effective countermeasures in areas where pests and vermin appear is a major challenge in terms of safety management of living environments and workplaces. Conventional systems often detect and counter these pests manually, making it difficult to respond quickly. Furthermore, there is a lack of technology to quickly implement optimal countermeasures for specific pests.

[0200] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0201] In this invention, the server includes means for collecting video and audio using an imaging device and an audio collecting device installed in an area where pests or vermin appear, means for transmitting the collected video and audio data to a mobile communication terminal, means for analyzing the data received by the mobile communication terminal and identifying targets, means for automatically generating countermeasures corresponding to the identified targets, means for transmitting instructions for executing the generated countermeasures, means for executing the countermeasures based on the received instructions, and means for re-monitoring and analyzing the situation after execution. This enables early detection of pests and rapid and effective countermeasures.

[0202] - "Imaging device" refers to hardware for capturing images.

[0203] "Sound collection device" refers to hardware for collecting sound.

[0204] "Mobile communication terminal" refers to a terminal that uses a mobile communication network, such as a mobile phone or smartphone.

[0205] "Analyzing the data" refers to the process of processing collected video and audio data to identify specific information.

[0206] "Target" refers to pests or vermin identified from the analyzed data.

[0207] "Automatically generating countermeasures" refers to the process by which the system automatically determines the optimal countermeasure based on the identified objectives.

[0208] "Directives" refer to the specific commands required to execute the generated countermeasures.

[0209] "Implementing countermeasures" refers to the process of actually taking countermeasures based on instructions.

[0210] "Re-monitoring and analyzing" refers to the process of collecting the situation after the countermeasures have been implemented and performing the same analysis.

[0211] The present invention relates to a system for early detection and rapid response in areas where pests or vermin appear. This system includes a series of processes that collect video and audio using an imaging device and an audio collecting device, transmit the video and audio to a mobile communication terminal for analysis, identify targets, and automatically generate and execute countermeasures.

[0212] Overall system overview

[0213] The system consists of the following major components:

[0214] 1. Imaging and audio collection devices: These are devices that are installed in the target area and continuously collect video and audio.

[0215] 2. Mobile communication terminal: Usually a smartphone, which receives the collected data and analyzes it using AI models.

[0216] 3. Server: Communicates with mobile communication terminals and sends commands.

[0217] Program processing procedure

[0218] The server first receives the video and audio transmitted from the imaging and audio collection devices. It then preprocesses the received data, removing noise and normalizing it. This process is performed using software such as OpenCV (for video processing) and Librosa (for audio processing).

[0219] The server then uses AI models to perform image and speech recognition, using deep learning libraries such as Tensorflow and Keras. For example, YOLO and Faster R-CNN are used for image recognition, and WaveNet and Transformer-based models are used for speech recognition.

[0220] Once a target is identified, the server automatically generates a countermeasure accordingly: for example, if a bear is identified, a countermeasure that plays a loud cracking sound is generated, and if a hornet is identified, a countermeasure that generates ultrasound is selected.

[0221] The generated countermeasures are sent as commands to the mobile communication device, which then executes the actual countermeasures based on the received commands, such as playing a sound from the smartphone speaker or emitting light or ultrasound from a specific device.

[0222] Specific examples

[0223] For example, if a bear appears near a house, the system operates as follows:

[0224] The imaging and sound collection devices capture images and sounds of the bear and transmit the data to a mobile communication terminal.

[0225] The mobile communication terminal receives the data and performs noise removal and normalization processing.

[0226] Next, an analysis is performed using an AI model to identify the bear.

[0227] Once a bear is identified, a command is sent from the server to the mobile communication terminal to play a loud cracking sound.

[0228] The mobile device will then play a clicking sound to scare off the bears, and will continue to monitor the area and take further action if necessary.

[0229] This system allows users to automatically detect pests and vermin early and take swift and effective measures to maintain a safe living and working environment.

[0230] Prompt Sentence Examples

[0231] "Please create a program that analyzes camera footage and audio from around the home to detect bears and hornets. Please also include a function that automatically suggests and executes appropriate countermeasures when specific pests or vermin are detected."

[0232] The above is an embodiment of the present invention.

[0233] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0234] Step 1: Gather information

[0235] The terminal uses an imaging device and an audio collecting device to capture video and audio within the area in real time.

[0236] Input: Raw video and audio data captured from imaging and audio collection devices.

[0237] Data processing: Format conversion and temporary storage of captured data.

[0238] Output: Format-converted and temporarily stored video and audio data.

[0239] Step 2: Send data

[0240] The terminal transmits the captured video and audio data to a server via a network.

[0241] Input: Format-converted and buffered video and audio data.

[0242] Data processing: Packetizing data and sending it over the network.

[0243] Output: Data packets sent to the server.

[0244] Step 3: Data reception and preprocessing

[0245] The server receives the data packets sent from the terminal and restores them as video and audio data, as well as performs noise removal and data normalization.

[0246] Input: Data packets sent from the terminal.

[0247] Data processing: Reformatting data packets, removing noise, and normalizing data.

[0248] Output: Pre-processed video and audio data.

[0249] Step 4: Data analysis and target identification

[0250] The server analyzes the pre-processed video and audio data using a generative AI model to identify targets, using deep learning models such as YOLO, Faster R-CNN, and WaveNet using Tensorflow and Keras.

[0251] Input: Preprocessed video and audio data.

[0252] Data computation: Deep learning analysis using AI models.

[0253] Output: Information about identified targets (bears, hornets, etc.).

[0254] Step 5: Generate solutions

[0255] The server automatically generates the best response depending on the identified target, for example, a loud cracking sound for a bear or ultrasonic waves for a hornet.

[0256] Input: Identified landmark information.

[0257] Data calculation: Execution of algorithm for automatic generation of countermeasures.

[0258] Output: Generated action instructions.

[0259] Step 6: Send instructions for how to deal with the problem

[0260] The server sends a command to the terminal to execute the generated countermeasure.

[0261] Input: Generated action instructions.

[0262] Data processing: Packetizing commands and sending them over the network.

[0263] Output: Command packets sent to the terminal.

[0264] Step 7: Implementing the solution

[0265] The device then takes action based on the received command, for example by playing a specific sound from a smartphone speaker or emitting light or ultrasound from a specific device.

[0266] Input: Received command packet.

[0267] Specific actions: playing sound, emitting light, generating ultrasound.

[0268] Output: The results of the action taken.

[0269] Step 8: Remonitor and reanalyze

[0270] After implementing the countermeasures, the device continues to monitor the area and transmits newly captured video and audio data to the server again. The server analyzes the data again and, if necessary, implements the countermeasures again.

[0271] Input: Newly captured video and audio data.

[0272] Data calculation: Pre-processing for reanalysis and analysis using AI models.

[0273] Specific actions: Continuous monitoring and repeated analysis within the area.

[0274] Output: Decision whether the situation is safe or whether further action is required.

[0275] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0276] ---

[0277] MODE FOR CARRYING OUT THE INVENTION

[0278] The present invention combines a system that uses a camera and a microphone to collect video and audio within an area, analyzes the video and audio to identify the target, and automatically executes an appropriate countermeasure, with an emotion engine that recognizes the user's emotions. This emotion engine makes it possible to optimize and improve the countermeasure, resulting in a more effective response. The following describes in detail the embodiments of the present invention.

[0279] Overall system overview

[0280] This system consists of cameras and microphones (hereinafter referred to as terminals) installed within the area, a server that receives and analyzes the data sent from these terminals and implements countermeasures, and an emotion engine that recognizes the user's emotions.

[0281] Program processing

[0282] 1. Information gathering

[0283] The device captures video and audio within the designated area in real time. This data is used as basic information to understand the situation within the area. The device also captures the user's facial expressions and comments at the same time.

[0284] 2. Data Transmission

[0285] The device transmits the captured video and audio data to a server via a network, enabling real-time data collection and transmission.

[0286] 3. Data Analysis

[0287] The server analyzes the received video and audio data, and performs preprocessing such as noise removal and data normalization.

[0288] The server uses an AI model to identify objects from video data and detect abnormal sounds from audio data.

[0289] At the same time, the server utilizes an emotion engine to recognize the user's emotions, which includes facial expression analysis and voice analysis.

[0290] 4. Generating solutions

[0291] The server automatically generates the most appropriate response to the object identified through the analysis: for example, a loud cracking sound is generated if a bear is identified, and ultrasonic waves are generated if a hornet is identified.

[0292] The content and strength of the coping strategies are optimized based on the user's recognized emotions. For example, if the user expresses fear, a stronger coping strategy is selected.

[0293] 5. Sending solutions

[0294] The server sends the generated command for the countermeasure to the terminal. This command includes the countermeasure to be executed and its parameters (e.g., type and intensity of sound, light pattern, etc.).

[0295] 6. Implementing the solution

[0296] The device will then take action based on the command it receives, for example, playing a repellent sound from a specific speaker if it receives a voice command, or emitting a bright light if it receives a video command.

[0297] 7. Recheck and Readdress

[0298] After the countermeasure is implemented, the device continues to monitor the situation within the area and captures video and audio data, which is also sent to the server again.

[0299] The server uses the emotion engine to analyze the transmitted data again and check how the user's emotions have changed.

[0300] Based on the analysis, it is determined whether additional measures are necessary. If the object is still present, appropriate measures are generated again and sent to the device. Repeat from step 4 as necessary.

[0301] Specific examples

[0302] For example, if a bear appears near a house, the system operates as follows:

[0303] The terminal (a camera installed around the house) captures video of the bear and sends that data along with the user's video and audio data to the server.

[0304] The server preprocesses the received video and uses an AI model to identify bears, while an emotion engine recognizes the user's emotions (e.g., fear or surprise).

[0305] The server generates a loud cracking sound that is effective against bears, and generates and sends to the terminal a command to adjust the volume according to the user's sense of fear.

[0306] The device will play a loud crackling sound to scare off the bears, and will continue to monitor the situation to see if any further action is needed.

[0307] Similarly, if a hornet's nest is found on the eaves of a house, the device captures video and audio, and the server analyzes them to identify the nest. The server generates ultrasonic waves that repel hornets and sends a command to the device to emit them at an intensity that corresponds to the user's emotions (e.g., anxiety). The device repels hornets by emitting ultrasonic waves.

[0308] This system allows users to detect the appearance of pests and vermin early and deal with them effectively, thereby providing a safe living environment. It also provides a sense of psychological security by optimizing countermeasures that take into account the user's emotions. This entire process is fully automated, so users do not need any special technical knowledge or operations to operate the system.

[0309] The above is a specific embodiment for carrying out the present invention.

[0310] The processing flow will be explained below.

[0311] Step-by-step process

[0312] Step 1:

[0313] The device captures video and audio within the set area in real time. Video is captured using a camera, and audio is collected using a microphone. The device also captures the user's facial expressions and voice. This data is temporarily stored in local storage and updated periodically.

[0314] Step 2:

[0315] The devices transmit the collected video and audio data over a network to a server. This data is streamed in real time or sent in batches. The data transmission is encrypted to ensure security.

[0316] Step 3:

[0317] The server receives the video and audio data sent from the device, saves the data in the data storage, and copies it to the workspace for analysis.

[0318] Step 4:

[0319] The server performs preprocessing on the received data, specifically noise removal, data normalization, and missing value imputation, to improve data quality and enhance analysis accuracy.

[0320] Step 5:

[0321] The server inputs the preprocessed data into the AI ​​model to detect objects. It uses an image recognition model to identify objects from video data and a voice recognition model to detect abnormal sounds from audio data.

[0322] Step 6:

[0323] The server uses an emotion engine to analyze the user's facial expressions and voice to recognize the user's emotions, such as surprise, fear, relief, etc.

[0324] Step 7:

[0325] The server automatically generates the most appropriate response to the object identified through the analysis: for example, a loud cracking sound is generated if a bear is identified, and ultrasonic waves are generated if a hornet is identified.

[0326] The content and strength of the coping strategies are optimized based on the user's recognized emotions. For example, if the user expresses fear, the system will adjust the frequency of the coping strategies.

[0327] Step 8:

[0328] The server sends the generated command for the countermeasure to the terminal. This command includes the countermeasure to be executed and its parameters (e.g., type and intensity of sound, light pattern, etc.).

[0329] Step 9:

[0330] The device will then take action based on the command it receives, for example, playing a repellent sound from a specific speaker if a voice command is received, or emitting a bright light if a video command is received.

[0331] Step 10:

[0332] The device will monitor the situation again after the countermeasure is implemented, capturing video and audio data, which will also be sent to the server again.

[0333] Step 11:

[0334] The server analyzes the retransmitted data and determines whether additional measures are necessary. If changes are observed in the object and the user's emotions have improved, the server determines that the measures are complete.

[0335] The server uses the emotion engine to analyze the transmitted data again and check how the user's emotions have changed.

[0336] If the server determines that further action is required, it generates an appropriate solution again and sends it to the terminal.

[0337] The above is the flow of program processing including specific operations.

[0338] Example 2

[0339] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0340] Conventional pest and vermin control systems have the problem of being unable to consider the psychological state of users, who have emotions, making it difficult to select the optimal countermeasure. Furthermore, conventional systems have low accuracy in identifying targets, making them prone to misidentification or failure to recognize them. This not only fails to adequately ensure user safety, but can also result in unnecessary countermeasures being taken. Furthermore, they lack the functionality to provide appropriate feedback on changes in the situation or the user's emotions after implementing countermeasures.

[0341] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recognizing the user's emotions, a means for automatically generating a countermeasure according to the identified object and the user's emotions, and a means for re-monitoring and analyzing the situation after execution and the user's emotions. This allows for the selection of an appropriate and effective countermeasure for the object, and also enables an optimal response that takes the user's emotions into consideration. As a result, the user's safety and psychological sense of security can be improved.

[0342] "Pests and vermin" refers to animals or insects that may appear in an area and cause human or physical harm.

[0343] "Camera" refers to a photographic device for capturing video data.

[0344] "Microphone" refers to an acoustic receiving device for capturing audio data.

[0345] "Collection means" refers to a combination of equipment and software for collecting video and audio data from within a specified area.

[0346] "Means for receiving" refers to the communications network and software for transmitting collected video and audio data to the server.

[0347] "Means for analyzing and identifying objects" refers to the processes and devices that use AI models and algorithms to identify specific pests or vermin from the data received.

[0348] "Means for recognizing user emotions" refers to a combination of software and hardware for analyzing the user's facial expressions and voice and determining their psychological state.

[0349] "Means for automatically generating countermeasures" refers to algorithms and programs for deriving optimal countermeasures based on the identified object and the user's emotions.

[0350] "Means for sending commands" refers to a communication network and software that sends instructions to the terminal to execute the generated countermeasure.

[0351] "Means for implementing countermeasures" refers to the equipment and software that implements the actual countermeasures (such as sound or light patterns) within the area based on the transmitted commands.

[0352] "Means for re-monitoring and re-analysis" refers to a combination of equipment and software for continuously monitoring the situation in an area after countermeasures have been implemented and for re-analyzing as necessary.

[0353] The present invention is a system that uses a camera and a microphone to collect and analyze video and audio within an area, identify the object, and automatically execute appropriate countermeasures. This system incorporates an emotion engine that recognizes the user's emotions, enabling optimization and improvement of countermeasures to achieve more effective responses. Detailed embodiments of the present invention are described below.

[0354] System Overview

[0355] This system consists of cameras and microphones (hereinafter referred to as terminals) installed within the area, a server that receives and analyzes the data sent from these terminals and implements countermeasures, and an emotion engine that recognizes the user's emotions.

[0356] Information gathering

[0357] The device captures video and audio within a set area in real time using a camera (e.g., Logitech C920) and a microphone (e.g., Blue Yeti). The collected data includes the objects within the area as well as the user's facial expressions and speech.

[0358] Data transmission

[0359] The device transmits the captured video and audio data to a server via a network (e.g., Wi-Fi or LAN), thereby achieving real-time data collection and transmission.

[0360] Data analysis

[0361] The server preprocesses the received video and audio data using OpenCV for video preprocessing and Librosa for audio preprocessing to remove noise and normalize the data.

[0362] The server uses an AI model (e.g., TensorFlow or PyTorch) to identify objects from video data and detect abnormal sounds from audio data.

[0363] At the same time, the server utilizes an emotion engine (e.g., Azure Face API) to recognize the user's emotions, which includes facial expression analysis and voice analysis.

[0364] Generate and send a workaround

[0365] The server generates an optimal response based on the identified object and the user's emotions, for example, if a bear is identified, a loud cracking sound will be generated and increased if the user expresses fear.

[0366] The server sends the generated command for the solution to the terminal, which includes the solution to be executed and its parameters.

[0367] Implementing the solution

[0368] The device will then take action based on the received command, for example, playing a repellent sound from a speaker (e.g., JBL Charge 3) or emitting a strong light using an LED light.

[0369] Recheck and re-address

[0370] Even after implementing a countermeasure, the device continues to monitor the situation within the area and captures video and audio. This data is then sent back to the server, which uses an emotion engine to check for changes in the user's emotions. If further action is required as a result of the analysis, an appropriate countermeasure is regenerated and sent to the device.

[0371] Specific examples

[0372] For example, if a bear appears near a house, the system operates as follows:

[0373] The terminals (cameras and microphones installed around the home) capture video and audio of the bear and send the data to a server.

[0374] The server preprocesses the received video and audio, and uses an AI model to identify bears. In parallel, an emotion engine recognizes the user's emotions (e.g., fear or surprise).

[0375] The server generates a loud cracking sound that is effective against bears, and generates and sends to the terminal a command to adjust the volume according to the user's sense of fear.

[0376] The device plays a loud crackling sound from the speaker (JBL Charge 3) to scare off bears, then continues to monitor the area to see if any further action is needed.

[0377] In this way, the system of the present invention can detect the appearance of pests and vermin at an early stage, ensuring the safety and psychological security of users.

[0378] Prompt Sentence Examples

[0379] "Please explain how the system works if a bear is in the vicinity of a residential area."

[0380] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0381] Step 1: Gather information

[0382] The device uses a camera and microphone to capture video and audio within a set area in real time. The input is the situation within the area and the user's facial expressions and comments, and the output is the captured video and audio data. Specifically, the camera detects movement within the area and collects video data, while the microphone collects audio within the area and generates audio data.

[0383] Step 2: Send data

[0384] The terminal sends the captured video and audio data to the server via the network. The input is the video and audio data obtained in step 1, and the output is the data sent to the server. Specifically, the terminal generates video and audio data packets and sends them to the server via the network. If an error occurs during data transmission, it attempts to resend them.

[0385] Step 3: Data analysis

[0386] The server analyzes the received video and audio data. First, as preprocessing, it uses OpenCV to remove noise and normalize the video data, and Librosa to remove noise and normalize the audio data. The input is the video and audio data sent in step 2, and the output is the preprocessed data. Specifically, it performs color adjustment and noise removal on the video data, and performs frequency analysis on the audio data.

[0387] Step 4: Identifying the Object and Emotion

[0388] The server uses the AI ​​model to identify objects from the preprocessed video data and uses an emotion engine to recognize the user's emotions from the audio and video data. The input is the data preprocessed in step 3, and the output is the identified objects and recognized emotion data. Specifically, the video data is input into the AI ​​model to identify the object (e.g., a bear), and the emotion engine analyzes the user's facial expressions and tone of voice to recognize their emotions.

[0389] Step 5: Generate solutions

[0390] The server generates an optimal response based on the identified object and the recognized user emotion. The input is the object and emotion data obtained in step 4, and the output is a command for the response. Specifically, it uses an AI algorithm to generate an effective response to the object (e.g., a loud crackling sound), and adjusts the parameters of the response (volume, duration, etc.) according to the user emotion.

[0391] Step 6: Submit your solution

[0392] The server sends the generated remedy command to the terminal. The input is the remedy command generated in step 5, and the output is the command sent to the terminal. Specifically, the remedy command is packetized and sent to the terminal via the network.

[0393] Step 7: Implementing the solution

[0394] The device executes countermeasures based on the received command. The input is the command sent from the server in step 6, and the output is the executed countermeasure. Specific actions include playing a specific repellent sound (e.g., a loud cracking sound) from the speaker and emitting a strong light from the LED light if necessary.

[0395] Step 8: Reassess and readdress

[0396] Even after the countermeasure is implemented, the device continues to monitor the situation within the area, capturing video and audio and sending it back to the server. The input is the new video and audio data after the countermeasure is implemented, and the output is the data sent back to the server. The server analyzes the data again and checks for changes in the user's emotions. The input is the data after the countermeasure is implemented, and the output is the results of the reanalysis and a new countermeasure, if necessary. If the analysis shows that further action is necessary, an appropriate countermeasure is regenerated and sent to the device.

[0397] (Application example 2)

[0398] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0399] In recent years, the importance of security in living and social environments has increased. However, existing security systems have limitations in their ability to detect intrusions by suspicious individuals and prevent the intrusion of pests and vermin. Furthermore, they are unable to respond flexibly to user emotional states, making it difficult to provide a sense of psychological security. Therefore, there is a need for a system that can replace conventional systems, monitor situations in real time, automatically generate appropriate countermeasures, and provide optimal countermeasures that take the user's emotional state into account.

[0400] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0401] In this invention, the server includes means for collecting video and audio using a camera and microphone, means for receiving the collected video and audio data, means for analyzing the received data and identifying the target, and means for recognizing the user's emotion and optimizing the content and strength of the countermeasure based on the emotion, thereby enabling more effective and flexible security measures.

[0402] A "camera" is a device for collecting video data.

[0403] A "microphone" is a device for collecting audio data.

[0404] The "collection means" is a means for acquiring video and audio using a camera and microphone.

[0405] The "receiving means" is a means for transmitting collected video and audio data to a server via a network.

[0406] The "analysis means" is a means used to analyze the received video and audio data and identify objects or abnormalities.

[0407] An "object" is a specific entity within an area, such as an obstacle, suspicious person, pest, or vermin identified by the analysis means.

[0408] The "automatic countermeasure generation means" is a means for automatically generating an optimal countermeasure for a specified object.

[0409] The "command transmission means" is a means for transmitting an instruction to execute the generated countermeasure to the terminal.

[0410] The "countermeasure execution means" is a means for executing a specific countermeasure based on a received command.

[0411] The "re-monitoring means" is a means for re-monitoring the situation after the countermeasure has been implemented and collecting video and audio data again.

[0412] "Emotion recognition means" refers to a means for analyzing and recognizing a user's emotions.

[0413] The "optimization method" is a method for optimizing the content and intensity of countermeasures based on the recognized user emotions.

[0414] This invention is a system that uses cameras and microphones to collect video and audio in areas where pests and vermin appear, analyzes the video and audio to identify obstacles and suspicious individuals, and automatically generates countermeasures. This system is combined with an emotion recognition engine that recognizes the user's emotions and optimizes the content and strength of the countermeasures.

[0415] System Program

[0416] The system consists of the following components:

[0417] 1. Terminal

[0418] It includes a camera for collecting video data and a microphone for collecting audio data, and the cameras and microphones are installed inside and outside the home.

[0419] 2. Server

[0420] Data receiving means: Receives video and audio data transmitted from the terminal.

[0421] Data analysis method: Analyzes received video and audio data to identify objects (vermin, vermin, intruders, etc.). Face detection and abnormal sound detection models using OpenCV are used for the analysis.

[0422] Emotion recognition: Recognizes the user's emotions from the collected data and optimizes the strength and content of countermeasures based on that data. For emotion recognition, an emotion recognition model using Keras is used.

[0423] Countermeasure generation means: Automatically generates the optimal countermeasure based on the analysis results. For example, if an intruder is detected, a command to play an acoustic alarm is generated.

[0424] Command sending means: Sends the generated instructions for the solution to the terminal.

[0425] 3. Users

[0426] The system collects the user's emotional data (e.g., fear, surprise, etc.) and takes appropriate action based on that emotion.

[0427] Hardware and software used

[0428] Camera: Hardware for collecting video footage.

[0429] Microphone: Hardware for collecting sound.

[0430] Server: A central processing unit for data analysis and emotion recognition. A server with Python, OpenCV, and Keras installed is used.

[0431] Acoustic alarm device: Used as one of the countermeasures.

[0432] Specific examples of processing

[0433] For example, if an intruder appears near a home, the system will:

[0434] The devices (cameras and microphones installed around the house) capture video and audio of intruders and send the data to a server.

[0435] The server preprocesses the received video and audio data and uses an AI model to identify intruders. In parallel, an emotion recognition engine recognizes the user's emotions (e.g., fear or surprise).

[0436] The server generates an effective acoustic alarm against an intruder, and generates and sends to the terminal a command to adjust the intensity of the alarm according to the user's emotion data.

[0437] The device will play an audible alarm to scare off intruders, then continue monitoring to see if any further action is required.

[0438] Examples of prompt statements

[0439] "Design a smartphone application that runs inside the home. It uses a camera and microphone to capture video and audio in real time to detect intruders and suspicious activity. It uses an emotion recognition engine to detect user emotions (especially fear and surprise), and if an abnormality is detected, it will sound an alarm and notify the user or security company. Using Python, OpenCV, and Keras, you will pre-train the emotion recognition model and perform real-time analysis."

[0440] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0441] Step 1:

[0442] The device uses a camera and microphone to collect video and audio data in real time. This data includes the situation within the area, the user's facial expressions, and their comments. The collected data serves as basic information for understanding the situation within the area.

[0443] Step 2:

[0444] The terminals transmit the collected video and audio data to a server via a network. This transmission is performed in real time, and it is important to minimize delays.

[0445] Step 3:

[0446] The server preprocesses the received video and audio data, specifically removing noise and normalizing the data. This preprocessing improves data quality and increases analysis accuracy.

[0447] Step 4:

[0448] The server analyzes the preprocessed data. It uses an AI model to identify objects (such as suspicious people or pests) from the video data, and detects abnormal sounds (such as the sound of destruction or metal) from the audio data. The analysis results are used to confirm the presence of the object.

[0449] Step 5:

[0450] The server uses an emotion recognition engine to analyze the user's emotions. By analyzing facial expressions from video data and changes in tone and volume from audio data, it can determine what emotions the user is feeling (for example, fear or surprise).

[0451] Step 6:

[0452] The server generates the optimal countermeasure based on the analysis results. Specifically, it automatically determines the best action to take against the target object (for example, playing a loud cracking sound or emitting ultrasound). It also adjusts the strength and content of the countermeasure according to the user's emotions.

[0453] Step 7:

[0454] The server sends the generated instructions (commands) for the countermeasures to the terminal. These commands include the countermeasures to be executed and their parameters (e.g., type and intensity of sound, light pattern, etc.).

[0455] Step 8:

[0456] The device will then take action based on the command it receives. For example, if it receives a voice command, it will play a repellent sound from a specific speaker, and if it receives a video command, it will emit a strong light. This will allow it to take action against the target.

[0457] Step 9:

[0458] Even after the countermeasure is implemented, the device continues to monitor the situation within the area, collecting video and audio data again and sending it to the server.

[0459] Step 10:

[0460] The server analyzes the data again to see how the user's emotions have changed. Based on the analysis results, it determines whether further action is necessary. If the object still exists and further action is necessary, the process repeats from step 6.

[0461] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0462] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0463] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0464] [Second embodiment]

[0465] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0466] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0467] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0468] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0469] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0470] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0471] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0472] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0473] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0474] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0475] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0476] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0477] ---

[0478] MODE FOR CARRYING OUT THE INVENTION

[0479] The present invention relates to a system for detecting the appearance of pests and vermin at an early stage and automatically taking appropriate measures. This system includes a series of processes that use a camera and a microphone to collect video and audio within an area, analyze the video and audio to identify the target, and take appropriate measures. The following describes in detail an embodiment of the present invention.

[0480] Overall system overview

[0481] This system consists of cameras and microphones (hereafter referred to as terminals) installed within the area, and a server that receives and analyzes the data sent from these terminals and implements countermeasures.

[0482] Program processing

[0483] 1. Information gathering

[0484] The device captures video and audio within the set area in real time, and this data is used as the basis for understanding the situation within the area.

[0485] 2. Data Transmission

[0486] The device transmits the captured video and audio data to a server via a network, enabling real-time data collection and transmission.

[0487] 3. Data Analysis

[0488] The server analyzes the received video and audio data, and performs preprocessing such as noise removal and data normalization.

[0489] The server uses an AI model (e.g., an image recognition model or a voice recognition model using deep learning) to identify the object (e.g., a bear or a hornet).

[0490] 4. Generating solutions

[0491] The server automatically generates the most effective response to the object identified through the analysis: for example, a loud cracking sound for a bear, or ultrasonic waves for a hornet.

[0492] 5. Sending solutions

[0493] The server sends the generated command for the solution to the terminal, which includes the specific solution and the timing for its execution.

[0494] 6. Implementing the solution

[0495] The device will then take action based on the command it receives, for example, playing a specific sound or emitting a bright light.

[0496] 7. Recheck and Readdress

[0497] After the countermeasure is implemented, the device continues to monitor the situation within the area and captures video and audio again, which are then sent to the server for further analysis.

[0498] If the server determines that further action is necessary, it generates an appropriate countermeasure and sends it to the device. By repeating this process, the target object is completely avoided.

[0499] Specific examples

[0500] For example, if a bear appears near a house, the system operates as follows:

[0501] The terminals (cameras installed around the home) capture images of bears and send the data to a server.

[0502] The server preprocesses the received footage and uses an AI model to identify the bear.

[0503] The server generates a command to play a loud cracking sound that is effective against bears and sends it to the device.

[0504] The device plays a loud cracking sound to scare off bears.

[0505] The device will capture the video again to check if the bear has completely left the area, and take further action if necessary.

[0506] Similarly, if a hornet's nest is discovered on the eaves of a house, the device captures video and audio, and the server analyzes them to identify the nest. The server selects a countermeasure to generate ultrasonic waves that repel hornets and sends it to the device. The device repels hornets by emitting ultrasonic waves.

[0507] This system allows users to detect the appearance of pests and vermin early and deal with them effectively, thereby providing a safe living environment. This series of processes is fully automated, so users do not need any special technical knowledge or operation to operate the system.

[0508] The above is a specific embodiment for carrying out the present invention.

[0509] The processing flow will be explained below.

[0510] Step-by-step process

[0511] Step 1:

[0512] The device captures video and audio within the set area in real time. Video is captured using a camera and audio is collected using a microphone. The collected data is temporarily stored in local storage and updated periodically.

[0513] Step 2:

[0514] The devices transmit the collected video and audio data over a network to a server. This data is streamed in real time or sent in batches. The data transmission is encrypted to ensure security.

[0515] Step 3:

[0516] The server receives the video and audio data sent from the device, saves the data in the data storage, and copies it to the workspace for analysis.

[0517] Step 4:

[0518] The server performs preprocessing on the received data, specifically noise removal, data normalization, and missing value imputation, to improve data quality and enhance analysis accuracy.

[0519] Step 5:

[0520] The server inputs the preprocessed data into the AI ​​model to detect objects. It uses an image recognition model to identify objects from video data and a voice recognition model to detect abnormal sounds from audio data.

[0521] Step 6:

[0522] The server automatically generates the most appropriate response to the object identified through the analysis: for example, a loud cracking sound is generated if a bear is identified, and ultrasonic waves are generated if a hornet is identified.

[0523] Step 7:

[0524] The server sends the generated command for the countermeasure to the terminal. This command includes the countermeasure to be executed and its parameters (e.g., type and intensity of sound, light pattern, etc.).

[0525] Step 8:

[0526] The device will then take action based on the command it receives, for example, playing a repellent sound from a specific speaker if a voice command is received, or emitting a bright light if a video command is received.

[0527] Step 9:

[0528] The device will monitor the situation again after the countermeasure is implemented, capturing video and audio data, which will also be sent to the server again.

[0529] Step 10:

[0530] The server analyzes the retransmitted data and determines whether additional action is required. If the object is still present, it generates an appropriate action again and sends it to the device. Repeat from step 6 if necessary.

[0531] The above is the flow of the program processing including the specific operations at each step.

[0532] Example 1

[0533] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0534] Conventional systems for detecting and responding to the emergence of pests and vermin have problems with real-time accuracy, often resulting in delayed responses and insufficient effectiveness. Furthermore, manual operation is required, and users without special technical knowledge have had the problem of being unable to operate the systems properly. Furthermore, the methods offered as avoidance measures by the systems are uniform, making it difficult to provide optimal countermeasures tailored to the characteristics of the target.

[0535] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0536] In this invention, the server includes a means for analyzing collected video and audio data to identify an object, a means for automatically generating a countermeasure corresponding to the identified object, and a means for transmitting a command for executing the generated countermeasure, thereby enabling high-precision detection of an object in real time and automatic appropriate countermeasures.

[0537] An "image capture device" is a device installed to capture video within a specific area in real time.

[0538] "Audio recording device" refers to equipment installed to record audio within a specific area.

[0539] "Means for collecting" refers to the function or process of acquiring video and audio data using image capture devices and audio recording devices.

[0540] The "means for receiving" is a function or process for transmitting collected video and audio data to a server via a network and for the server to receive it.

[0541] "Means for analyzing" refers to algorithms or systems that process received video and audio data and identify objects.

[0542] The "means for automatic generation" is a function or process that automatically creates an appropriate countermeasure for the target object based on the results of data analysis.

[0543] The "means for sending a command" is a function or process that sends an instruction to execute the generated countermeasure to the terminal via the network.

[0544] The "means for executing" refers to a function or device that the terminal actually physically executes the countermeasure based on the received command.

[0545] "Means for monitoring and analysis" refers to the function or process of rechecking the situation after the countermeasures have been implemented and reanalyzing it based on newly collected data.

[0546] The "pre-processing means" refers to a function or process that performs noise removal and data normalization on the collected video and audio data.

[0547] A "generative AI model" is a model or algorithm that uses deep learning or machine learning to identify objects from data.

[0548] MODE FOR CARRYING OUT THE INVENTION

[0549] The present invention provides a system for detecting the appearance of pests and vermin at an early stage and automatically taking appropriate countermeasures. This system includes a series of processes that use an image capture device and an audio recording device to collect video and audio within an area, analyze the video and audio, identify the target, and take appropriate countermeasures. The following describes in detail the embodiments of the present invention.

[0550] Overall system overview

[0551] This system consists of image capture devices and audio recording devices (hereafter referred to as terminals) installed within the area, and a server that receives and analyzes the data sent from these terminals and executes countermeasures. The system is designed to be easy to operate, even if the user does not have special technical knowledge.

[0552] Program processing

[0553] 1. Information collection via devices

[0554] The device captures video and audio within the set area in real time. The image capture device captures video at 30 frames per second, and the audio recorder records audio at 24-bit / 48kHz.

[0555] 2. Data transmission by the terminal

[0556] The device periodically compresses the captured video and audio data, converts it into a data structure such as JSON or Protobuf format, and sends it to the server using HTTP / HTTPS or WebSocket protocols.

[0557] 3. Data analysis by the server (preprocessing)

[0558] The server first decodes the received data and performs noise reduction on the video and audio. The video data is decoded using the OpenCV library to reduce blur and noise, and the audio data is decoded using Python's pydub and SciPy.

[0559] 4. Data analysis by the server (identification of the target object)

[0560] The server analyzes the video and audio data using a generative AI model (for example, a deep learning model using TensorFlow or PyTorch) and uses object detection algorithms such as YOLO (You Only Look Once) or R-CNN to identify objects such as bears or hornets.

[0561] 5. Server-generated solutions

[0562] Based on the identified object, the server uses a pre-defined rules-based system to generate the optimal response, such as playing a loud cracking sound for a bear or generating ultrasound for a hornet.

[0563] 6. Server sends solution

[0564] The server sends the generated command for the solution to the appropriate device at the appropriate time. This command includes the specific content of the solution (e.g., the path of the audio file, playback timing, etc.).

[0565] 7. Implementing measures on the device

[0566] The device will then take action based on the command it receives, such as playing a loud crackling sound through the speaker or flashing an LED light.

[0567] 8. Reconfirmation and re-action by terminal

[0568] After implementing a countermeasure, the device continues to monitor the situation within the area and recaptures video and audio data. This data is then sent to the server for further analysis. If the server determines that further countermeasures are necessary, it generates a new appropriate countermeasure and sends it to the device. This process is repeated as many times as necessary.

[0569] Specific examples

[0570] For example, if a bear appears near a house, the system operates as follows:

[0571] 1. Information collection via devices

[0572] The device captures video at 30 frames per second using an image capture device installed in the garden, and records audio at 24-bit / 48kHz using an audio recording device.

[0573] 2. Data transmission by the terminal

[0574] The device compresses the captured data every 5 seconds, converts it into JSON format, and sends it to the server using the HTTP protocol.

[0575] 3. Data analysis by the server (preprocessing)

[0576] The server uses OpenCV to remove noise from the received video data and pydub to perform noise filtering on the audio data.

[0577] 4. Data analysis by the server (identification of the target object)

[0578] The server uses the YOLO model to analyze the video data and identify the presence of bears.

[0579] 5. Server-generated solutions

[0580] The server generates commands based on pre-defined rules to play a loud cracking sound that works on bears.

[0581] 6. Server sends solution

[0582] The server sends the generated command (e.g., the path to the audio file and the playback timing) to the terminal.

[0583] 7. Implementing measures on the device

[0584] Based on the command received, the device plays a loud cracking sound through its speaker to scare off bears.

[0585] 8. Reconfirmation and re-action by terminal

[0586] The device captures the video again to check if the bear has completely left the area. If the bear is still in the area, the server generates a new countermeasure and sends it to the device to execute again.

[0587] Example prompts for generative AI models

[0588] "Please tell me how to use this system to automatically detect the appearance of bears and take appropriate action."

[0589] This is a specific mode for carrying out the invention.

[0590] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0591] Program processing steps

[0592] Step 1: Gather information

[0593] The device captures video and audio within the area in real time. Specifically, the image capture device captures video at 30 frames per second, and the audio recording device records audio at 24-bit / 48kHz.

[0594] Input: Physical conditions within the set area

[0595] Output: Video and audio data

[0596] Step 2: Send data

[0597] The device periodically compresses the captured video and audio data, converts it into JSON or Protobuf format, and sends it to the server using HTTP / HTTPS or WebSocket protocol. For example, the device compresses and sends the data every 5 seconds.

[0598] Input: Video and audio data

[0599] Output: Compressed data

[0600] Step 3: Data analysis (preprocessing)

[0601] The server decodes the received data and performs noise reduction on the video and audio. Specifically, the OpenCV library is used to reduce blur and noise in the video data, and pydub and SciPy are used to perform noise filtering on the audio data.

[0602] Input: Compressed data

[0603] Output: Denoised data

[0604] Step 4: Data analysis (target identification)

[0605] The server analyzes the video and audio data using a deep learning model (e.g., a generative AI model using TensorFlow or PyTorch) to identify objects using object detection algorithms such as YOLO or R-CNN.

[0606] Input: Denoised data

[0607] Output: Information about the identified object (e.g., location and type of bear or hornet)

[0608] Step 5: Generate solutions

[0609] Based on the identified object, the server uses a rules engine to automatically generate the optimal response, for example, a loud cracking sound for a bear or ultrasonic waves for a hornet.

[0610] Input: Information about the identified object

[0611] Output: Action command

[0612] Step 6: Submit your solution

[0613] The server generates a command for the appropriate action and sends it to the device at the appropriate time, which includes specific instructions such as the path to the audio file and the timing of playback.

[0614] Input: Solution command

[0615] Output: Commands sent to the terminal

[0616] Step 7: Implementing the solution

[0617] The device will then take action based on the command it receives, such as playing a loud crackling sound through the speaker or flashing an LED light.

[0618] Input: Command sent to the terminal

[0619] Output: Actions taken

[0620] Step 8: Reassess and readdress

[0621] Even after the device has implemented the countermeasures, it continues to monitor the situation within the area, recapturing video and audio data and sending it to the server. If it determines that reanalysis is necessary, it generates an appropriate countermeasure again and sends it to the device.

[0622] Input: Post-execution video and audio data

[0623] Output: Updated workaround command

[0624] The above is the processing flow of the specific program of the system.

[0625] (Application example 1)

[0626] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0627] Automating the early detection of pests and effective countermeasures in areas where pests and vermin appear is a major challenge in terms of safety management of living environments and workplaces. Conventional systems often detect and counter these pests manually, making it difficult to respond quickly. Furthermore, there is a lack of technology to quickly implement optimal countermeasures for specific pests.

[0628] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0629] In this invention, the server includes means for collecting video and audio using an imaging device and an audio collecting device installed in an area where pests or vermin appear, means for transmitting the collected video and audio data to a mobile communication terminal, means for analyzing the data received by the mobile communication terminal and identifying targets, means for automatically generating countermeasures corresponding to the identified targets, means for transmitting instructions for executing the generated countermeasures, means for executing the countermeasures based on the received instructions, and means for re-monitoring and analyzing the situation after execution. This enables early detection of pests and rapid and effective countermeasures.

[0630] - "Imaging device" refers to hardware for capturing images.

[0631] "Sound collection device" refers to hardware for collecting sound.

[0632] "Mobile communication terminal" refers to a terminal that uses a mobile communication network, such as a mobile phone or smartphone.

[0633] "Analyzing the data" refers to the process of processing collected video and audio data to identify specific information.

[0634] "Target" refers to pests or vermin identified from the analyzed data.

[0635] "Automatically generating countermeasures" refers to the process by which the system automatically determines the optimal countermeasure based on the identified objectives.

[0636] "Directives" refer to the specific commands required to execute the generated countermeasures.

[0637] "Implementing countermeasures" refers to the process of actually taking countermeasures based on instructions.

[0638] "Re-monitoring and analyzing" refers to the process of collecting the situation after the countermeasures have been implemented and performing the same analysis.

[0639] The present invention relates to a system for early detection and rapid response in areas where pests or vermin appear. This system includes a series of processes that collect video and audio using an imaging device and an audio collecting device, transmit the video and audio to a mobile communication terminal for analysis, identify targets, and automatically generate and execute countermeasures.

[0640] Overall system overview

[0641] The system consists of the following major components:

[0642] 1. Imaging and audio collection devices: These are devices that are installed in the target area and continuously collect video and audio.

[0643] 2. Mobile communication terminal: Usually a smartphone, which receives the collected data and analyzes it using AI models.

[0644] 3. Server: Communicates with mobile communication terminals and sends commands.

[0645] Program processing procedure

[0646] The server first receives the video and audio transmitted from the imaging and audio collection devices. It then preprocesses the received data, removing noise and normalizing it. This process is performed using software such as OpenCV (for video processing) and Librosa (for audio processing).

[0647] The server then uses AI models to perform image and speech recognition, using deep learning libraries such as Tensorflow and Keras. For example, YOLO and Faster R-CNN are used for image recognition, and WaveNet and Transformer-based models are used for speech recognition.

[0648] Once a target is identified, the server automatically generates a countermeasure accordingly: for example, if a bear is identified, a countermeasure that plays a loud cracking sound is generated, and if a hornet is identified, a countermeasure that generates ultrasound is selected.

[0649] The generated countermeasures are sent as commands to the mobile communication device, which then executes the actual countermeasures based on the received commands, such as playing a sound from the smartphone speaker or emitting light or ultrasound from a specific device.

[0650] Specific examples

[0651] For example, if a bear appears near a house, the system operates as follows:

[0652] The imaging and sound collection devices capture images and sounds of the bear and transmit the data to a mobile communication terminal.

[0653] The mobile communication terminal receives the data and performs noise removal and normalization processing.

[0654] Next, an analysis is performed using an AI model to identify the bear.

[0655] Once a bear is identified, a command is sent from the server to the mobile communication terminal to play a loud cracking sound.

[0656] The mobile device will then play a clicking sound to scare off the bears, and will continue to monitor the area and take further action if necessary.

[0657] This system allows users to automatically detect pests and vermin early and take swift and effective measures to maintain a safe living and working environment.

[0658] Prompt Sentence Examples

[0659] "Please create a program that analyzes camera footage and audio from around the home to detect bears and hornets. Please also include a function that automatically suggests and executes appropriate countermeasures when specific pests or vermin are detected."

[0660] The above is an embodiment of the present invention.

[0661] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0662] Step 1: Gather information

[0663] The terminal uses an imaging device and an audio collecting device to capture video and audio within the area in real time.

[0664] Input: Raw video and audio data captured from imaging and audio collection devices.

[0665] Data processing: Format conversion and temporary storage of captured data.

[0666] Output: Format-converted and temporarily stored video and audio data.

[0667] Step 2: Send data

[0668] The terminal transmits the captured video and audio data to a server via a network.

[0669] Input: Format-converted and buffered video and audio data.

[0670] Data processing: Packetizing data and sending it over the network.

[0671] Output: Data packets sent to the server.

[0672] Step 3: Data reception and preprocessing

[0673] The server receives the data packets sent from the terminal and restores them as video and audio data, as well as performs noise removal and data normalization.

[0674] Input: Data packets sent from the terminal.

[0675] Data processing: Reformatting data packets, removing noise, and normalizing data.

[0676] Output: Pre-processed video and audio data.

[0677] Step 4: Data analysis and target identification

[0678] The server analyzes the pre-processed video and audio data using a generative AI model to identify targets, using deep learning models such as YOLO, Faster R-CNN, and WaveNet using Tensorflow and Keras.

[0679] Input: Preprocessed video and audio data.

[0680] Data computation: Deep learning analysis using AI models.

[0681] Output: Information about identified targets (bears, hornets, etc.).

[0682] Step 5: Generate solutions

[0683] The server automatically generates the best response depending on the identified target, for example, a loud cracking sound for a bear or ultrasonic waves for a hornet.

[0684] Input: Identified landmark information.

[0685] Data calculation: Execution of algorithm for automatic generation of countermeasures.

[0686] Output: Generated action instructions.

[0687] Step 6: Send instructions for how to deal with the problem

[0688] The server sends a command to the terminal to execute the generated countermeasure.

[0689] Input: Generated action instructions.

[0690] Data processing: Packetizing commands and sending them over the network.

[0691] Output: Command packets sent to the terminal.

[0692] Step 7: Implementing the solution

[0693] The device then takes action based on the received command, for example by playing a specific sound from a smartphone speaker or emitting light or ultrasound from a specific device.

[0694] Input: Received command packet.

[0695] Specific actions: playing sound, emitting light, generating ultrasound.

[0696] Output: The results of the action taken.

[0697] Step 8: Remonitor and reanalyze

[0698] After implementing the countermeasures, the device continues to monitor the area and transmits newly captured video and audio data to the server again. The server analyzes the data again and, if necessary, implements the countermeasures again.

[0699] Input: Newly captured video and audio data.

[0700] Data calculation: Pre-processing for reanalysis and analysis using AI models.

[0701] Specific actions: Continuous monitoring and repeated analysis within the area.

[0702] Output: Decision whether the situation is safe or whether further action is required.

[0703] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0704] ---

[0705] MODE FOR CARRYING OUT THE INVENTION

[0706] The present invention combines a system that uses a camera and a microphone to collect video and audio within an area, analyzes the video and audio to identify the target, and automatically executes an appropriate countermeasure, with an emotion engine that recognizes the user's emotions. This emotion engine makes it possible to optimize and improve the countermeasure, resulting in a more effective response. The following describes in detail the embodiments of the present invention.

[0707] Overall system overview

[0708] This system consists of cameras and microphones (hereinafter referred to as terminals) installed within the area, a server that receives and analyzes the data sent from these terminals and implements countermeasures, and an emotion engine that recognizes the user's emotions.

[0709] Program processing

[0710] 1. Information gathering

[0711] The device captures video and audio within the designated area in real time. This data is used as basic information to understand the situation within the area. The device also captures the user's facial expressions and comments at the same time.

[0712] 2. Data Transmission

[0713] The device transmits the captured video and audio data to a server via a network, enabling real-time data collection and transmission.

[0714] 3. Data Analysis

[0715] The server analyzes the received video and audio data, and performs preprocessing such as noise removal and data normalization.

[0716] The server uses an AI model to identify objects from video data and detect abnormal sounds from audio data.

[0717] At the same time, the server utilizes an emotion engine to recognize the user's emotions, which includes facial expression analysis and voice analysis.

[0718] 4. Generating solutions

[0719] The server automatically generates the most appropriate response to the object identified through the analysis: for example, a loud cracking sound is generated if a bear is identified, and ultrasonic waves are generated if a hornet is identified.

[0720] The content and strength of the coping strategies are optimized based on the user's recognized emotions. For example, if the user expresses fear, a stronger coping strategy is selected.

[0721] 5. Sending solutions

[0722] The server sends the generated command for the countermeasure to the terminal. This command includes the countermeasure to be executed and its parameters (e.g., type and intensity of sound, light pattern, etc.).

[0723] 6. Implementing the solution

[0724] The device will then take action based on the command it receives, for example, playing a repellent sound from a specific speaker if it receives a voice command, or emitting a bright light if it receives a video command.

[0725] 7. Recheck and Readdress

[0726] After the countermeasure is implemented, the device continues to monitor the situation within the area and captures video and audio data, which is also sent to the server again.

[0727] The server uses the emotion engine to analyze the transmitted data again and check how the user's emotions have changed.

[0728] Based on the analysis, it is determined whether additional measures are necessary. If the object is still present, appropriate measures are generated again and sent to the device. Repeat from step 4 as necessary.

[0729] Specific examples

[0730] For example, if a bear appears near a house, the system operates as follows:

[0731] The terminal (a camera installed around the house) captures video of the bear and sends that data along with the user's video and audio data to the server.

[0732] The server preprocesses the received video and uses an AI model to identify bears, while an emotion engine recognizes the user's emotions (e.g., fear or surprise).

[0733] The server generates a loud cracking sound that is effective against bears, and generates and sends to the terminal a command to adjust the volume according to the user's sense of fear.

[0734] The device will play a loud crackling sound to scare off the bears, and will continue to monitor the situation to see if any further action is needed.

[0735] Similarly, if a hornet's nest is found on the eaves of a house, the device captures video and audio, and the server analyzes them to identify the nest. The server generates ultrasonic waves that repel hornets and sends a command to the device to emit them at an intensity that corresponds to the user's emotions (e.g., anxiety). The device repels hornets by emitting ultrasonic waves.

[0736] This system allows users to detect the appearance of pests and vermin early and deal with them effectively, thereby providing a safe living environment. It also provides a sense of psychological security by optimizing countermeasures that take into account the user's emotions. This entire process is fully automated, so users do not need any special technical knowledge or operations to operate the system.

[0737] The above is a specific embodiment for carrying out the present invention.

[0738] The processing flow will be explained below.

[0739] Step-by-step process

[0740] Step 1:

[0741] The device captures video and audio within the set area in real time. Video is captured using a camera, and audio is collected using a microphone. The device also captures the user's facial expressions and voice. This data is temporarily stored in local storage and updated periodically.

[0742] Step 2:

[0743] The devices transmit the collected video and audio data over a network to a server. This data is streamed in real time or sent in batches. The data transmission is encrypted to ensure security.

[0744] Step 3:

[0745] The server receives the video and audio data sent from the device, saves the data in the data storage, and copies it to the workspace for analysis.

[0746] Step 4:

[0747] The server performs preprocessing on the received data, specifically noise removal, data normalization, and missing value imputation, to improve data quality and enhance analysis accuracy.

[0748] Step 5:

[0749] The server inputs the preprocessed data into the AI ​​model to detect objects. It uses an image recognition model to identify objects from video data and a voice recognition model to detect abnormal sounds from audio data.

[0750] Step 6:

[0751] The server uses an emotion engine to analyze the user's facial expressions and voice to recognize the user's emotions, such as surprise, fear, relief, etc.

[0752] Step 7:

[0753] The server automatically generates the most appropriate response to the object identified through the analysis: for example, a loud cracking sound is generated if a bear is identified, and ultrasonic waves are generated if a hornet is identified.

[0754] The content and strength of the coping strategies are optimized based on the user's recognized emotions. For example, if the user expresses fear, the system will adjust the frequency of the coping strategies.

[0755] Step 8:

[0756] The server sends the generated command for the countermeasure to the terminal. This command includes the countermeasure to be executed and its parameters (e.g., type and intensity of sound, light pattern, etc.).

[0757] Step 9:

[0758] The device will then take action based on the command it receives, for example, playing a repellent sound from a specific speaker if a voice command is received, or emitting a bright light if a video command is received.

[0759] Step 10:

[0760] The device will monitor the situation again after the countermeasure is implemented, capturing video and audio data, which will also be sent to the server again.

[0761] Step 11:

[0762] The server analyzes the retransmitted data and determines whether additional measures are necessary. If changes are observed in the object and the user's emotions have improved, the server determines that the measures are complete.

[0763] The server uses the emotion engine to analyze the transmitted data again and check how the user's emotions have changed.

[0764] If the server determines that further action is required, it generates an appropriate solution again and sends it to the terminal.

[0765] The above is the flow of program processing including specific operations.

[0766] Example 2

[0767] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0768] Conventional pest and vermin control systems have the problem of being unable to consider the psychological state of users, who have emotions, making it difficult to select the optimal countermeasure. Furthermore, conventional systems have low accuracy in identifying targets, making them prone to misidentification or failure to recognize them. This not only fails to adequately ensure user safety, but can also result in unnecessary countermeasures being taken. Furthermore, they lack the functionality to provide appropriate feedback on changes in the situation or the user's emotions after implementing countermeasures.

[0769] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recognizing the user's emotions, a means for automatically generating a countermeasure according to the identified object and the user's emotions, and a means for re-monitoring and analyzing the situation after execution and the user's emotions. This allows for the selection of an appropriate and effective countermeasure for the object, and also enables an optimal response that takes the user's emotions into consideration. As a result, the user's safety and psychological sense of security can be improved.

[0770] "Pests and vermin" refers to animals or insects that may appear in an area and cause human or physical harm.

[0771] "Camera" refers to a photographic device for capturing video data.

[0772] "Microphone" refers to an acoustic receiving device for capturing audio data.

[0773] "Collection means" refers to a combination of equipment and software for collecting video and audio data from within a specified area.

[0774] "Means for receiving" refers to the communications network and software for transmitting collected video and audio data to the server.

[0775] "Means for analyzing and identifying objects" refers to the processes and devices that use AI models and algorithms to identify specific pests or vermin from the data received.

[0776] "Means for recognizing user emotions" refers to a combination of software and hardware for analyzing the user's facial expressions and voice and determining their psychological state.

[0777] "Means for automatically generating countermeasures" refers to algorithms and programs for deriving optimal countermeasures based on the identified object and the user's emotions.

[0778] "Means for sending commands" refers to a communication network and software that sends instructions to the terminal to execute the generated countermeasure.

[0779] "Means for implementing countermeasures" refers to the equipment and software that implements the actual countermeasures (such as sound or light patterns) within the area based on the transmitted commands.

[0780] "Means for re-monitoring and re-analysis" refers to a combination of equipment and software for continuously monitoring the situation in an area after countermeasures have been implemented and for re-analyzing as necessary.

[0781] The present invention is a system that uses a camera and a microphone to collect and analyze video and audio within an area, identify the object, and automatically execute appropriate countermeasures. This system incorporates an emotion engine that recognizes the user's emotions, enabling optimization and improvement of countermeasures to achieve more effective responses. Detailed embodiments of the present invention are described below.

[0782] System Overview

[0783] This system consists of cameras and microphones (hereinafter referred to as terminals) installed within the area, a server that receives and analyzes the data sent from these terminals and implements countermeasures, and an emotion engine that recognizes the user's emotions.

[0784] Information gathering

[0785] The device captures video and audio within a set area in real time using a camera (e.g., Logitech C920) and a microphone (e.g., Blue Yeti). The collected data includes the objects within the area as well as the user's facial expressions and speech.

[0786] Data transmission

[0787] The device transmits the captured video and audio data to a server via a network (e.g., Wi-Fi or LAN), thereby achieving real-time data collection and transmission.

[0788] Data analysis

[0789] The server preprocesses the received video and audio data using OpenCV for video preprocessing and Librosa for audio preprocessing to remove noise and normalize the data.

[0790] The server uses an AI model (e.g., TensorFlow or PyTorch) to identify objects from video data and detect abnormal sounds from audio data.

[0791] At the same time, the server utilizes an emotion engine (e.g., Azure Face API) to recognize the user's emotions, which includes facial expression analysis and voice analysis.

[0792] Generate and send a workaround

[0793] The server generates an optimal response based on the identified object and the user's emotions, for example, if a bear is identified, a loud cracking sound will be generated and increased if the user expresses fear.

[0794] The server sends the generated command for the solution to the terminal, which includes the solution to be executed and its parameters.

[0795] Implementing the solution

[0796] The device will then take action based on the received command, for example, playing a repellent sound from a speaker (e.g., JBL Charge 3) or emitting a strong light using an LED light.

[0797] Recheck and re-address

[0798] Even after implementing a countermeasure, the device continues to monitor the situation within the area and captures video and audio. This data is then sent back to the server, which uses an emotion engine to check for changes in the user's emotions. If further action is required as a result of the analysis, an appropriate countermeasure is regenerated and sent to the device.

[0799] Specific examples

[0800] For example, if a bear appears near a house, the system operates as follows:

[0801] The terminals (cameras and microphones installed around the home) capture video and audio of the bear and send the data to a server.

[0802] The server preprocesses the received video and audio, and uses an AI model to identify bears. In parallel, an emotion engine recognizes the user's emotions (e.g., fear or surprise).

[0803] The server generates a loud cracking sound that is effective against bears, and generates and sends to the terminal a command to adjust the volume according to the user's sense of fear.

[0804] The device plays a loud crackling sound from the speaker (JBL Charge 3) to scare off bears, then continues to monitor the area to see if any further action is needed.

[0805] In this way, the system of the present invention can detect the appearance of pests and vermin at an early stage, ensuring the safety and psychological security of users.

[0806] Prompt Sentence Examples

[0807] "Please explain how the system works if a bear is in the vicinity of a residential area."

[0808] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0809] Step 1: Gather information

[0810] The device uses a camera and microphone to capture video and audio within a set area in real time. The input is the situation within the area and the user's facial expressions and comments, and the output is the captured video and audio data. Specifically, the camera detects movement within the area and collects video data, while the microphone collects audio within the area and generates audio data.

[0811] Step 2: Send data

[0812] The terminal sends the captured video and audio data to the server via the network. The input is the video and audio data obtained in step 1, and the output is the data sent to the server. Specifically, the terminal generates video and audio data packets and sends them to the server via the network. If an error occurs during data transmission, it attempts to resend them.

[0813] Step 3: Data analysis

[0814] The server analyzes the received video and audio data. First, as preprocessing, it uses OpenCV to remove noise and normalize the video data, and Librosa to remove noise and normalize the audio data. The input is the video and audio data sent in step 2, and the output is the preprocessed data. Specifically, it performs color adjustment and noise removal on the video data, and performs frequency analysis on the audio data.

[0815] Step 4: Identifying the Object and Emotion

[0816] The server uses the AI ​​model to identify objects from the preprocessed video data and uses an emotion engine to recognize the user's emotions from the audio and video data. The input is the data preprocessed in step 3, and the output is the identified objects and recognized emotion data. Specifically, the video data is input into the AI ​​model to identify the object (e.g., a bear), and the emotion engine analyzes the user's facial expressions and tone of voice to recognize their emotions.

[0817] Step 5: Generate solutions

[0818] The server generates an optimal response based on the identified object and the recognized user emotion. The input is the object and emotion data obtained in step 4, and the output is a command for the response. Specifically, it uses an AI algorithm to generate an effective response to the object (e.g., a loud crackling sound), and adjusts the parameters of the response (volume, duration, etc.) according to the user emotion.

[0819] Step 6: Submit your solution

[0820] The server sends the generated remedy command to the terminal. The input is the remedy command generated in step 5, and the output is the command sent to the terminal. Specifically, the remedy command is packetized and sent to the terminal via the network.

[0821] Step 7: Implementing the solution

[0822] The device executes countermeasures based on the received command. The input is the command sent from the server in step 6, and the output is the executed countermeasure. Specific actions include playing a specific repellent sound (e.g., a loud cracking sound) from the speaker and emitting a strong light from the LED light if necessary.

[0823] Step 8: Reassess and readdress

[0824] Even after the countermeasure is implemented, the device continues to monitor the situation within the area, capturing video and audio and sending it back to the server. The input is the new video and audio data after the countermeasure is implemented, and the output is the data sent back to the server. The server analyzes the data again and checks for changes in the user's emotions. The input is the data after the countermeasure is implemented, and the output is the results of the reanalysis and a new countermeasure, if necessary. If the analysis shows that further action is necessary, an appropriate countermeasure is regenerated and sent to the device.

[0825] (Application example 2)

[0826] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0827] In recent years, the importance of security in living and social environments has increased. However, existing security systems have limitations in their ability to detect intrusions by suspicious individuals and prevent the intrusion of pests and vermin. Furthermore, they are unable to respond flexibly to user emotional states, making it difficult to provide a sense of psychological security. Therefore, there is a need for a system that can replace conventional systems, monitor situations in real time, automatically generate appropriate countermeasures, and provide optimal countermeasures that take the user's emotional state into account.

[0828] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0829] In this invention, the server includes means for collecting video and audio using a camera and microphone, means for receiving the collected video and audio data, means for analyzing the received data and identifying the target, and means for recognizing the user's emotion and optimizing the content and strength of the countermeasure based on the emotion, thereby enabling more effective and flexible security measures.

[0830] A "camera" is a device for collecting video data.

[0831] A "microphone" is a device for collecting audio data.

[0832] The "collection means" is a means for acquiring video and audio using a camera and microphone.

[0833] The "receiving means" is a means for transmitting collected video and audio data to a server via a network.

[0834] The "analysis means" is a means used to analyze the received video and audio data and identify objects or abnormalities.

[0835] An "object" is a specific entity within an area, such as an obstacle, suspicious person, pest, or vermin identified by the analysis means.

[0836] The "automatic countermeasure generation means" is a means for automatically generating an optimal countermeasure for a specified object.

[0837] The "command transmission means" is a means for transmitting an instruction to execute the generated countermeasure to the terminal.

[0838] The "countermeasure execution means" is a means for executing a specific countermeasure based on a received command.

[0839] The "re-monitoring means" is a means for re-monitoring the situation after the countermeasure has been implemented and collecting video and audio data again.

[0840] "Emotion recognition means" refers to a means for analyzing and recognizing a user's emotions.

[0841] The "optimization method" is a method for optimizing the content and intensity of countermeasures based on the recognized user emotions.

[0842] This invention is a system that uses cameras and microphones to collect video and audio in areas where pests and vermin appear, analyzes the video and audio to identify obstacles and suspicious individuals, and automatically generates countermeasures. This system is combined with an emotion recognition engine that recognizes the user's emotions and optimizes the content and strength of the countermeasures.

[0843] System Program

[0844] The system consists of the following components:

[0845] 1. Terminal

[0846] It includes a camera for collecting video data and a microphone for collecting audio data, and the cameras and microphones are installed inside and outside the home.

[0847] 2. Server

[0848] Data receiving means: Receives video and audio data transmitted from the terminal.

[0849] Data analysis method: Analyzes received video and audio data to identify objects (vermin, vermin, intruders, etc.). Face detection and abnormal sound detection models using OpenCV are used for the analysis.

[0850] Emotion recognition: Recognizes the user's emotions from the collected data and optimizes the strength and content of countermeasures based on that data. For emotion recognition, an emotion recognition model using Keras is used.

[0851] Countermeasure generation means: Automatically generates the optimal countermeasure based on the analysis results. For example, if an intruder is detected, a command to play an acoustic alarm is generated.

[0852] Command sending means: Sends the generated instructions for the solution to the terminal.

[0853] 3. Users

[0854] The system collects the user's emotional data (e.g., fear, surprise, etc.) and takes appropriate action based on that emotion.

[0855] Hardware and software used

[0856] Camera: Hardware for collecting video footage.

[0857] Microphone: Hardware for collecting sound.

[0858] Server: A central processing unit for data analysis and emotion recognition. A server with Python, OpenCV, and Keras installed is used.

[0859] Acoustic alarm device: Used as one of the countermeasures.

[0860] Specific examples of processing

[0861] For example, if an intruder appears near a home, the system will:

[0862] The devices (cameras and microphones installed around the house) capture video and audio of intruders and send the data to a server.

[0863] The server preprocesses the received video and audio data and uses an AI model to identify intruders. In parallel, an emotion recognition engine recognizes the user's emotions (e.g., fear or surprise).

[0864] The server generates an effective acoustic alarm against an intruder, and generates and sends to the terminal a command to adjust the intensity of the alarm according to the user's emotion data.

[0865] The device will play an audible alarm to scare off intruders, then continue monitoring to see if any further action is required.

[0866] Examples of prompt statements

[0867] "Design a smartphone application that runs inside the home. It uses a camera and microphone to capture video and audio in real time to detect intruders and suspicious activity. It uses an emotion recognition engine to detect user emotions (especially fear and surprise), and if an abnormality is detected, it will sound an alarm and notify the user or security company. Using Python, OpenCV, and Keras, you will pre-train the emotion recognition model and perform real-time analysis."

[0868] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0869] Step 1:

[0870] The device uses a camera and microphone to collect video and audio data in real time. This data includes the situation within the area, the user's facial expressions, and their comments. The collected data serves as basic information for understanding the situation within the area.

[0871] Step 2:

[0872] The terminals transmit the collected video and audio data to a server via a network. This transmission is performed in real time, and it is important to minimize delays.

[0873] Step 3:

[0874] The server preprocesses the received video and audio data, specifically removing noise and normalizing the data. This preprocessing improves data quality and increases analysis accuracy.

[0875] Step 4:

[0876] The server analyzes the preprocessed data. It uses an AI model to identify objects (such as suspicious people or pests) from the video data, and detects abnormal sounds (such as the sound of destruction or metal) from the audio data. The analysis results are used to confirm the presence of the object.

[0877] Step 5:

[0878] The server uses an emotion recognition engine to analyze the user's emotions. By analyzing facial expressions from video data and changes in tone and volume from audio data, it can determine what emotions the user is feeling (for example, fear or surprise).

[0879] Step 6:

[0880] The server generates the optimal countermeasure based on the analysis results. Specifically, it automatically determines the best action to take against the target object (for example, playing a loud cracking sound or emitting ultrasound). It also adjusts the strength and content of the countermeasure according to the user's emotions.

[0881] Step 7:

[0882] The server sends the generated instructions (commands) for the countermeasures to the terminal. These commands include the countermeasures to be executed and their parameters (e.g., type and intensity of sound, light pattern, etc.).

[0883] Step 8:

[0884] The device will then take action based on the command it receives. For example, if it receives a voice command, it will play a repellent sound from a specific speaker, and if it receives a video command, it will emit a strong light. This will allow it to take action against the target.

[0885] Step 9:

[0886] Even after the countermeasure is implemented, the device continues to monitor the situation within the area, collecting video and audio data again and sending it to the server.

[0887] Step 10:

[0888] The server analyzes the data again to see how the user's emotions have changed. Based on the analysis results, it determines whether further action is necessary. If the object still exists and further action is necessary, the process repeats from step 6.

[0889] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0890] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0891] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0892] [Third embodiment]

[0893] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0894] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0895] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0896] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0897] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0898] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0899] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0900] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0901] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0902] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0903] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0904] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0905] ---

[0906] MODE FOR CARRYING OUT THE INVENTION

[0907] The present invention relates to a system for detecting the appearance of pests and vermin at an early stage and automatically taking appropriate measures. This system includes a series of processes that use a camera and a microphone to collect video and audio within an area, analyze the video and audio to identify the target, and take appropriate measures. The following describes in detail an embodiment of the present invention.

[0908] Overall system overview

[0909] This system consists of cameras and microphones (hereafter referred to as terminals) installed within the area, and a server that receives and analyzes the data sent from these terminals and implements countermeasures.

[0910] Program processing

[0911] 1. Information gathering

[0912] The device captures video and audio within the set area in real time, and this data is used as the basis for understanding the situation within the area.

[0913] 2. Data Transmission

[0914] The device transmits the captured video and audio data to a server via a network, enabling real-time data collection and transmission.

[0915] 3. Data Analysis

[0916] The server analyzes the received video and audio data, and performs preprocessing such as noise removal and data normalization.

[0917] The server uses an AI model (e.g., an image recognition model or a voice recognition model using deep learning) to identify the object (e.g., a bear or a hornet).

[0918] 4. Generating solutions

[0919] The server automatically generates the most effective response to the object identified through the analysis: for example, a loud cracking sound for a bear, or ultrasonic waves for a hornet.

[0920] 5. Sending solutions

[0921] The server sends the generated command for the solution to the terminal, which includes the specific solution and the timing for its execution.

[0922] 6. Implementing the solution

[0923] The device will then take action based on the command it receives, for example, playing a specific sound or emitting a bright light.

[0924] 7. Recheck and Readdress

[0925] After the countermeasure is implemented, the device continues to monitor the situation within the area and captures video and audio again, which are then sent to the server for further analysis.

[0926] If the server determines that further action is necessary, it generates an appropriate countermeasure and sends it to the device. By repeating this process, the target object is completely avoided.

[0927] Specific examples

[0928] For example, if a bear appears near a house, the system operates as follows:

[0929] The terminals (cameras installed around the home) capture images of bears and send the data to a server.

[0930] The server preprocesses the received footage and uses an AI model to identify the bear.

[0931] The server generates a command to play a loud cracking sound that is effective against bears and sends it to the device.

[0932] The device plays a loud cracking sound to scare off bears.

[0933] The device will capture the video again to check if the bear has completely left the area, and take further action if necessary.

[0934] Similarly, if a hornet's nest is discovered on the eaves of a house, the device captures video and audio, and the server analyzes them to identify the nest. The server selects a countermeasure to generate ultrasonic waves that repel hornets and sends it to the device. The device repels hornets by emitting ultrasonic waves.

[0935] This system allows users to detect the appearance of pests and vermin early and deal with them effectively, thereby providing a safe living environment. This series of processes is fully automated, so users do not need any special technical knowledge or operation to operate the system.

[0936] The above is a specific embodiment for carrying out the present invention.

[0937] The processing flow will be explained below.

[0938] Step-by-step process

[0939] Step 1:

[0940] The device captures video and audio within the set area in real time. Video is captured using a camera and audio is collected using a microphone. The collected data is temporarily stored in local storage and updated periodically.

[0941] Step 2:

[0942] The devices transmit the collected video and audio data over a network to a server. This data is streamed in real time or sent in batches. The data transmission is encrypted to ensure security.

[0943] Step 3:

[0944] The server receives the video and audio data sent from the device, saves the data in the data storage, and copies it to the workspace for analysis.

[0945] Step 4:

[0946] The server performs preprocessing on the received data, specifically noise removal, data normalization, and missing value imputation, to improve data quality and enhance analysis accuracy.

[0947] Step 5:

[0948] The server inputs the preprocessed data into the AI ​​model to detect objects. It uses an image recognition model to identify objects from video data and a voice recognition model to detect abnormal sounds from audio data.

[0949] Step 6:

[0950] The server automatically generates the most appropriate response to the object identified through the analysis: for example, a loud cracking sound is generated if a bear is identified, and ultrasonic waves are generated if a hornet is identified.

[0951] Step 7:

[0952] The server sends the generated command for the countermeasure to the terminal. This command includes the countermeasure to be executed and its parameters (e.g., type and intensity of sound, light pattern, etc.).

[0953] Step 8:

[0954] The device will then take action based on the command it receives, for example, playing a repellent sound from a specific speaker if a voice command is received, or emitting a bright light if a video command is received.

[0955] Step 9:

[0956] The device will monitor the situation again after the countermeasure is implemented, capturing video and audio data, which will also be sent to the server again.

[0957] Step 10:

[0958] The server analyzes the retransmitted data and determines whether additional action is required. If the object is still present, it generates an appropriate action again and sends it to the device. Repeat from step 6 if necessary.

[0959] The above is the flow of the program processing including the specific operations at each step.

[0960] Example 1

[0961] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0962] Conventional systems for detecting and responding to the emergence of pests and vermin have problems with real-time accuracy, often resulting in delayed responses and insufficient effectiveness. Furthermore, manual operation is required, and users without special technical knowledge have had the problem of being unable to operate the systems properly. Furthermore, the methods offered as avoidance measures by the systems are uniform, making it difficult to provide optimal countermeasures tailored to the characteristics of the target.

[0963] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0964] In this invention, the server includes a means for analyzing collected video and audio data to identify an object, a means for automatically generating a countermeasure corresponding to the identified object, and a means for transmitting a command for executing the generated countermeasure, thereby enabling high-precision detection of an object in real time and automatic appropriate countermeasures.

[0965] An "image capture device" is a device installed to capture video within a specific area in real time.

[0966] "Audio recording device" refers to equipment installed to record audio within a specific area.

[0967] "Means for collecting" refers to the function or process of acquiring video and audio data using image capture devices and audio recording devices.

[0968] The "means for receiving" is a function or process for transmitting collected video and audio data to a server via a network and for the server to receive it.

[0969] "Means for analyzing" refers to algorithms or systems that process received video and audio data and identify objects.

[0970] The "means for automatic generation" is a function or process that automatically creates an appropriate countermeasure for the target object based on the results of data analysis.

[0971] The "means for sending a command" is a function or process that sends an instruction to execute the generated countermeasure to the terminal via the network.

[0972] The "means for executing" refers to a function or device that the terminal actually physically executes the countermeasure based on the received command.

[0973] "Means for monitoring and analysis" refers to the function or process of rechecking the situation after the countermeasures have been implemented and reanalyzing it based on newly collected data.

[0974] The "pre-processing means" refers to a function or process that performs noise removal and data normalization on the collected video and audio data.

[0975] A "generative AI model" is a model or algorithm that uses deep learning or machine learning to identify objects from data.

[0976] MODE FOR CARRYING OUT THE INVENTION

[0977] The present invention provides a system for detecting the appearance of pests and vermin at an early stage and automatically taking appropriate countermeasures. This system includes a series of processes that use an image capture device and an audio recording device to collect video and audio within an area, analyze the video and audio, identify the target, and take appropriate countermeasures. The following describes in detail the embodiments of the present invention.

[0978] Overall system overview

[0979] This system consists of image capture devices and audio recording devices (hereafter referred to as terminals) installed within the area, and a server that receives and analyzes the data sent from these terminals and executes countermeasures. The system is designed to be easy to operate, even if the user does not have special technical knowledge.

[0980] Program processing

[0981] 1. Information collection via devices

[0982] The device captures video and audio within the set area in real time. The image capture device captures video at 30 frames per second, and the audio recorder records audio at 24-bit / 48kHz.

[0983] 2. Data transmission by the terminal

[0984] The device periodically compresses the captured video and audio data, converts it into a data structure such as JSON or Protobuf format, and sends it to the server using HTTP / HTTPS or WebSocket protocols.

[0985] 3. Data analysis by the server (preprocessing)

[0986] The server first decodes the received data and performs noise reduction on the video and audio. The video data is decoded using the OpenCV library to reduce blur and noise, and the audio data is decoded using Python's pydub and SciPy.

[0987] 4. Data analysis by the server (identification of the target object)

[0988] The server analyzes the video and audio data using a generative AI model (for example, a deep learning model using TensorFlow or PyTorch) and uses object detection algorithms such as YOLO (You Only Look Once) or R-CNN to identify objects such as bears or hornets.

[0989] 5. Server-generated solutions

[0990] Based on the identified object, the server uses a pre-defined rules-based system to generate the optimal response, such as playing a loud cracking sound for a bear or generating ultrasound for a hornet.

[0991] 6. Server sends solution

[0992] The server sends the generated command for the solution to the appropriate device at the appropriate time. This command includes the specific content of the solution (e.g., the path of the audio file, playback timing, etc.).

[0993] 7. Implementing measures on the device

[0994] The device will then take action based on the command it receives, such as playing a loud crackling sound through the speaker or flashing an LED light.

[0995] 8. Reconfirmation and re-action by terminal

[0996] After implementing a countermeasure, the device continues to monitor the situation within the area and recaptures video and audio data. This data is then sent to the server for further analysis. If the server determines that further countermeasures are necessary, it generates a new appropriate countermeasure and sends it to the device. This process is repeated as many times as necessary.

[0997] Specific examples

[0998] For example, if a bear appears near a house, the system operates as follows:

[0999] 1. Information collection via devices

[1000] The device captures video at 30 frames per second using an image capture device installed in the garden, and records audio at 24-bit / 48kHz using an audio recording device.

[1001] 2. Data transmission by the terminal

[1002] The device compresses the captured data every 5 seconds, converts it into JSON format, and sends it to the server using the HTTP protocol.

[1003] 3. Data analysis by the server (preprocessing)

[1004] The server uses OpenCV to remove noise from the received video data and pydub to perform noise filtering on the audio data.

[1005] 4. Data analysis by the server (identification of the target object)

[1006] The server uses the YOLO model to analyze the video data and identify the presence of bears.

[1007] 5. Server-generated solutions

[1008] The server generates commands based on pre-defined rules to play a loud cracking sound that works on bears.

[1009] 6. Server sends solution

[1010] The server sends the generated command (e.g., the path to the audio file and the playback timing) to the terminal.

[1011] 7. Implementing measures on the device

[1012] Based on the command received, the device plays a loud cracking sound through its speaker to scare off bears.

[1013] 8. Reconfirmation and re-action by terminal

[1014] The device captures the video again to check if the bear has completely left the area. If the bear is still in the area, the server generates a new countermeasure and sends it to the device to execute again.

[1015] Example prompts for generative AI models

[1016] "Please tell me how to use this system to automatically detect the appearance of bears and take appropriate action."

[1017] This is a specific mode for carrying out the invention.

[1018] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1019] Program processing steps

[1020] Step 1: Gather information

[1021] The device captures video and audio within the area in real time. Specifically, the image capture device captures video at 30 frames per second, and the audio recording device records audio at 24-bit / 48kHz.

[1022] Input: Physical conditions within the set area

[1023] Output: Video and audio data

[1024] Step 2: Send data

[1025] The device periodically compresses the captured video and audio data, converts it into JSON or Protobuf format, and sends it to the server using HTTP / HTTPS or WebSocket protocol. For example, the device compresses and sends the data every 5 seconds.

[1026] Input: Video and audio data

[1027] Output: Compressed data

[1028] Step 3: Data analysis (preprocessing)

[1029] The server decodes the received data and performs noise reduction on the video and audio. Specifically, the OpenCV library is used to reduce blur and noise in the video data, and pydub and SciPy are used to perform noise filtering on the audio data.

[1030] Input: Compressed data

[1031] Output: Denoised data

[1032] Step 4: Data analysis (target identification)

[1033] The server analyzes the video and audio data using a deep learning model (e.g., a generative AI model using TensorFlow or PyTorch) to identify objects using object detection algorithms such as YOLO or R-CNN.

[1034] Input: Denoised data

[1035] Output: Information about the identified object (e.g., location and type of bear or hornet)

[1036] Step 5: Generate solutions

[1037] Based on the identified object, the server uses a rules engine to automatically generate the optimal response, for example, a loud cracking sound for a bear or ultrasonic waves for a hornet.

[1038] Input: Information about the identified object

[1039] Output: Action command

[1040] Step 6: Submit your solution

[1041] The server generates a command for the appropriate action and sends it to the device at the appropriate time, which includes specific instructions such as the path to the audio file and the timing of playback.

[1042] Input: Solution command

[1043] Output: Commands sent to the terminal

[1044] Step 7: Implementing the solution

[1045] The device will then take action based on the command it receives, such as playing a loud crackling sound through the speaker or flashing an LED light.

[1046] Input: Command sent to the terminal

[1047] Output: Actions taken

[1048] Step 8: Reassess and readdress

[1049] Even after the device has implemented the countermeasures, it continues to monitor the situation within the area, recapturing video and audio data and sending it to the server. If it determines that reanalysis is necessary, it generates an appropriate countermeasure again and sends it to the device.

[1050] Input: Post-execution video and audio data

[1051] Output: Updated workaround command

[1052] The above is the processing flow of the specific program of the system.

[1053] (Application example 1)

[1054] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1055] Automating the early detection of pests and effective countermeasures in areas where pests and vermin appear is a major challenge in terms of safety management of living environments and workplaces. Conventional systems often detect and counter these pests manually, making it difficult to respond quickly. Furthermore, there is a lack of technology to quickly implement optimal countermeasures for specific pests.

[1056] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1057] In this invention, the server includes means for collecting video and audio using an imaging device and an audio collecting device installed in an area where pests or vermin appear, means for transmitting the collected video and audio data to a mobile communication terminal, means for analyzing the data received by the mobile communication terminal and identifying targets, means for automatically generating countermeasures corresponding to the identified targets, means for transmitting instructions for executing the generated countermeasures, means for executing the countermeasures based on the received instructions, and means for re-monitoring and analyzing the situation after execution. This enables early detection of pests and rapid and effective countermeasures.

[1058] - "Imaging device" refers to hardware for capturing images.

[1059] "Sound collection device" refers to hardware for collecting sound.

[1060] "Mobile communication terminal" refers to a terminal that uses a mobile communication network, such as a mobile phone or smartphone.

[1061] "Analyzing the data" refers to the process of processing collected video and audio data to identify specific information.

[1062] "Target" refers to pests or vermin identified from the analyzed data.

[1063] "Automatically generating countermeasures" refers to the process by which the system automatically determines the optimal countermeasure based on the identified objectives.

[1064] "Directives" refer to the specific commands required to execute the generated countermeasures.

[1065] "Implementing countermeasures" refers to the process of actually taking countermeasures based on instructions.

[1066] "Re-monitoring and analyzing" refers to the process of collecting the situation after the countermeasures have been implemented and performing the same analysis.

[1067] The present invention relates to a system for early detection and rapid response in areas where pests or vermin appear. This system includes a series of processes that collect video and audio using an imaging device and an audio collecting device, transmit the video and audio to a mobile communication terminal for analysis, identify targets, and automatically generate and execute countermeasures.

[1068] Overall system overview

[1069] The system consists of the following major components:

[1070] 1. Imaging and audio collection devices: These are devices that are installed in the target area and continuously collect video and audio.

[1071] 2. Mobile communication terminal: Usually a smartphone, which receives the collected data and analyzes it using AI models.

[1072] 3. Server: Communicates with mobile communication terminals and sends commands.

[1073] Program processing procedure

[1074] The server first receives the video and audio transmitted from the imaging and audio collection devices. It then preprocesses the received data, removing noise and normalizing it. This process is performed using software such as OpenCV (for video processing) and Librosa (for audio processing).

[1075] The server then uses AI models to perform image and speech recognition, using deep learning libraries such as Tensorflow and Keras. For example, YOLO and Faster R-CNN are used for image recognition, and WaveNet and Transformer-based models are used for speech recognition.

[1076] Once a target is identified, the server automatically generates a countermeasure accordingly: for example, if a bear is identified, a countermeasure that plays a loud cracking sound is generated, and if a hornet is identified, a countermeasure that generates ultrasound is selected.

[1077] The generated countermeasures are sent as commands to the mobile communication device, which then executes the actual countermeasures based on the received commands, such as playing a sound from the smartphone speaker or emitting light or ultrasound from a specific device.

[1078] Specific examples

[1079] For example, if a bear appears near a house, the system operates as follows:

[1080] The imaging and sound collection devices capture images and sounds of the bear and transmit the data to a mobile communication terminal.

[1081] The mobile communication terminal receives the data and performs noise removal and normalization processing.

[1082] Next, an analysis is performed using an AI model to identify the bear.

[1083] Once a bear is identified, a command is sent from the server to the mobile communication terminal to play a loud cracking sound.

[1084] The mobile device will then play a clicking sound to scare off the bears, and will continue to monitor the area and take further action if necessary.

[1085] This system allows users to automatically detect pests and vermin early and take swift and effective measures to maintain a safe living and working environment.

[1086] Prompt Sentence Examples

[1087] "Please create a program that analyzes camera footage and audio from around the home to detect bears and hornets. Please also include a function that automatically suggests and executes appropriate countermeasures when specific pests or vermin are detected."

[1088] The above is an embodiment of the present invention.

[1089] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1090] Step 1: Gather information

[1091] The terminal uses an imaging device and an audio collecting device to capture video and audio within the area in real time.

[1092] Input: Raw video and audio data captured from imaging and audio collection devices.

[1093] Data processing: Format conversion and temporary storage of captured data.

[1094] Output: Format-converted and temporarily stored video and audio data.

[1095] Step 2: Send data

[1096] The terminal transmits the captured video and audio data to a server via a network.

[1097] Input: Format-converted and buffered video and audio data.

[1098] Data processing: Packetizing data and sending it over the network.

[1099] Output: Data packets sent to the server.

[1100] Step 3: Data reception and preprocessing

[1101] The server receives the data packets sent from the terminal and restores them as video and audio data, as well as performs noise removal and data normalization.

[1102] Input: Data packets sent from the terminal.

[1103] Data processing: Reformatting data packets, removing noise, and normalizing data.

[1104] Output: Pre-processed video and audio data.

[1105] Step 4: Data analysis and target identification

[1106] The server analyzes the pre-processed video and audio data using a generative AI model to identify targets, using deep learning models such as YOLO, Faster R-CNN, and WaveNet using Tensorflow and Keras.

[1107] Input: Preprocessed video and audio data.

[1108] Data computation: Deep learning analysis using AI models.

[1109] Output: Information about identified targets (bears, hornets, etc.).

[1110] Step 5: Generate solutions

[1111] The server automatically generates the best response depending on the identified target, for example, a loud cracking sound for a bear or ultrasonic waves for a hornet.

[1112] Input: Identified landmark information.

[1113] Data calculation: Execution of algorithm for automatic generation of countermeasures.

[1114] Output: Generated action instructions.

[1115] Step 6: Send instructions for how to deal with the problem

[1116] The server sends a command to the terminal to execute the generated countermeasure.

[1117] Input: Generated action instructions.

[1118] Data processing: Packetizing commands and sending them over the network.

[1119] Output: Command packets sent to the terminal.

[1120] Step 7: Implementing the solution

[1121] The device then takes action based on the received command, for example by playing a specific sound from a smartphone speaker or emitting light or ultrasound from a specific device.

[1122] Input: Received command packet.

[1123] Specific actions: playing sound, emitting light, generating ultrasound.

[1124] Output: The results of the action taken.

[1125] Step 8: Remonitor and reanalyze

[1126] After implementing the countermeasures, the device continues to monitor the area and transmits newly captured video and audio data to the server again. The server analyzes the data again and, if necessary, implements the countermeasures again.

[1127] Input: Newly captured video and audio data.

[1128] Data calculation: Pre-processing for reanalysis and analysis using AI models.

[1129] Specific actions: Continuous monitoring and repeated analysis within the area.

[1130] Output: Decision whether the situation is safe or whether further action is required.

[1131] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1132] ---

[1133] MODE FOR CARRYING OUT THE INVENTION

[1134] The present invention combines a system that uses a camera and a microphone to collect video and audio within an area, analyzes the video and audio to identify the target, and automatically executes an appropriate countermeasure, with an emotion engine that recognizes the user's emotions. This emotion engine makes it possible to optimize and improve the countermeasure, resulting in a more effective response. The following describes in detail the embodiments of the present invention.

[1135] Overall system overview

[1136] This system consists of cameras and microphones (hereinafter referred to as terminals) installed within the area, a server that receives and analyzes the data sent from these terminals and implements countermeasures, and an emotion engine that recognizes the user's emotions.

[1137] Program processing

[1138] 1. Information gathering

[1139] The device captures video and audio within the designated area in real time. This data is used as basic information to understand the situation within the area. The device also captures the user's facial expressions and comments at the same time.

[1140] 2. Data Transmission

[1141] The device transmits the captured video and audio data to a server via a network, enabling real-time data collection and transmission.

[1142] 3. Data Analysis

[1143] The server analyzes the received video and audio data, and performs preprocessing such as noise removal and data normalization.

[1144] The server uses an AI model to identify objects from video data and detect abnormal sounds from audio data.

[1145] At the same time, the server utilizes an emotion engine to recognize the user's emotions, which includes facial expression analysis and voice analysis.

[1146] 4. Generating solutions

[1147] The server automatically generates the most appropriate response to the object identified through the analysis: for example, a loud cracking sound is generated if a bear is identified, and ultrasonic waves are generated if a hornet is identified.

[1148] The content and strength of the coping strategies are optimized based on the user's recognized emotions. For example, if the user expresses fear, a stronger coping strategy is selected.

[1149] 5. Sending solutions

[1150] The server sends the generated command for the countermeasure to the terminal. This command includes the countermeasure to be executed and its parameters (e.g., type and intensity of sound, light pattern, etc.).

[1151] 6. Implementing the solution

[1152] The device will then take action based on the command it receives, for example, playing a repellent sound from a specific speaker if it receives a voice command, or emitting a bright light if it receives a video command.

[1153] 7. Recheck and Readdress

[1154] After the countermeasure is implemented, the device continues to monitor the situation within the area and captures video and audio data, which is also sent to the server again.

[1155] The server uses the emotion engine to analyze the transmitted data again and check how the user's emotions have changed.

[1156] Based on the analysis, it is determined whether additional measures are necessary. If the object is still present, appropriate measures are generated again and sent to the device. Repeat from step 4 as necessary.

[1157] Specific examples

[1158] For example, if a bear appears near a house, the system operates as follows:

[1159] The terminal (a camera installed around the house) captures video of the bear and sends that data along with the user's video and audio data to the server.

[1160] The server preprocesses the received video and uses an AI model to identify bears, while an emotion engine recognizes the user's emotions (e.g., fear or surprise).

[1161] The server generates a loud cracking sound that is effective against bears, and generates and sends to the terminal a command to adjust the volume according to the user's sense of fear.

[1162] The device will play a loud crackling sound to scare off the bears, and will continue to monitor the situation to see if any further action is needed.

[1163] Similarly, if a hornet's nest is found on the eaves of a house, the device captures video and audio, and the server analyzes them to identify the nest. The server generates ultrasonic waves that repel hornets and sends a command to the device to emit them at an intensity that corresponds to the user's emotions (e.g., anxiety). The device repels hornets by emitting ultrasonic waves.

[1164] This system allows users to detect the appearance of pests and vermin early and deal with them effectively, thereby providing a safe living environment. It also provides a sense of psychological security by optimizing countermeasures that take into account the user's emotions. This entire process is fully automated, so users do not need any special technical knowledge or operations to operate the system.

[1165] The above is a specific embodiment for carrying out the present invention.

[1166] The processing flow will be explained below.

[1167] Step-by-step process

[1168] Step 1:

[1169] The device captures video and audio within the set area in real time. Video is captured using a camera, and audio is collected using a microphone. The device also captures the user's facial expressions and voice. This data is temporarily stored in local storage and updated periodically.

[1170] Step 2:

[1171] The devices transmit the collected video and audio data over a network to a server. This data is streamed in real time or sent in batches. The data transmission is encrypted to ensure security.

[1172] Step 3:

[1173] The server receives the video and audio data sent from the device, saves the data in the data storage, and copies it to the workspace for analysis.

[1174] Step 4:

[1175] The server performs preprocessing on the received data, specifically noise removal, data normalization, and missing value imputation, to improve data quality and enhance analysis accuracy.

[1176] Step 5:

[1177] The server inputs the preprocessed data into the AI ​​model to detect objects. It uses an image recognition model to identify objects from video data and a voice recognition model to detect abnormal sounds from audio data.

[1178] Step 6:

[1179] The server uses an emotion engine to analyze the user's facial expressions and voice to recognize the user's emotions, such as surprise, fear, relief, etc.

[1180] Step 7:

[1181] The server automatically generates the most appropriate response to the object identified through the analysis: for example, a loud cracking sound is generated if a bear is identified, and ultrasonic waves are generated if a hornet is identified.

[1182] The content and strength of the coping strategies are optimized based on the user's recognized emotions. For example, if the user expresses fear, the system will adjust the frequency of the coping strategies.

[1183] Step 8:

[1184] The server sends the generated command for the countermeasure to the terminal. This command includes the countermeasure to be executed and its parameters (e.g., type and intensity of sound, light pattern, etc.).

[1185] Step 9:

[1186] The device will then take action based on the command it receives, for example, playing a repellent sound from a specific speaker if a voice command is received, or emitting a bright light if a video command is received.

[1187] Step 10:

[1188] The device will monitor the situation again after the countermeasure is implemented, capturing video and audio data, which will also be sent to the server again.

[1189] Step 11:

[1190] The server analyzes the retransmitted data and determines whether additional measures are necessary. If changes are observed in the object and the user's emotions have improved, the server determines that the measures are complete.

[1191] The server uses the emotion engine to analyze the transmitted data again and check how the user's emotions have changed.

[1192] If the server determines that further action is required, it generates an appropriate solution again and sends it to the terminal.

[1193] The above is the flow of program processing including specific operations.

[1194] Example 2

[1195] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1196] Conventional pest and vermin control systems have the problem of being unable to consider the psychological state of users, who have emotions, making it difficult to select the optimal countermeasure. Furthermore, conventional systems have low accuracy in identifying targets, making them prone to misidentification or failure to recognize them. This not only fails to adequately ensure user safety, but can also result in unnecessary countermeasures being taken. Furthermore, they lack the functionality to provide appropriate feedback on changes in the situation or the user's emotions after implementing countermeasures.

[1197] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recognizing the user's emotions, a means for automatically generating a countermeasure according to the identified object and the user's emotions, and a means for re-monitoring and analyzing the situation after execution and the user's emotions. This allows for the selection of an appropriate and effective countermeasure for the object, and also enables an optimal response that takes the user's emotions into consideration. As a result, the user's safety and psychological sense of security can be improved.

[1198] "Pests and vermin" refers to animals or insects that may appear in an area and cause human or physical harm.

[1199] "Camera" refers to a photographic device for capturing video data.

[1200] "Microphone" refers to an acoustic receiving device for capturing audio data.

[1201] "Collection means" refers to a combination of equipment and software for collecting video and audio data from within a specified area.

[1202] "Means for receiving" refers to the communications network and software for transmitting collected video and audio data to the server.

[1203] "Means for analyzing and identifying objects" refers to the processes and devices that use AI models and algorithms to identify specific pests or vermin from the data received.

[1204] "Means for recognizing user emotions" refers to a combination of software and hardware for analyzing the user's facial expressions and voice and determining their psychological state.

[1205] "Means for automatically generating countermeasures" refers to algorithms and programs for deriving optimal countermeasures based on the identified object and the user's emotions.

[1206] "Means for sending commands" refers to a communication network and software that sends instructions to the terminal to execute the generated countermeasure.

[1207] "Means for implementing countermeasures" refers to the equipment and software that implements the actual countermeasures (such as sound or light patterns) within the area based on the transmitted commands.

[1208] "Means for re-monitoring and re-analysis" refers to a combination of equipment and software for continuously monitoring the situation in an area after countermeasures have been implemented and for re-analyzing as necessary.

[1209] The present invention is a system that uses a camera and a microphone to collect and analyze video and audio within an area, identify the object, and automatically execute appropriate countermeasures. This system incorporates an emotion engine that recognizes the user's emotions, enabling optimization and improvement of countermeasures to achieve more effective responses. Detailed embodiments of the present invention are described below.

[1210] System Overview

[1211] This system consists of cameras and microphones (hereinafter referred to as terminals) installed within the area, a server that receives and analyzes the data sent from these terminals and implements countermeasures, and an emotion engine that recognizes the user's emotions.

[1212] Information gathering

[1213] The device captures video and audio within a set area in real time using a camera (e.g., Logitech C920) and a microphone (e.g., Blue Yeti). The collected data includes the objects within the area as well as the user's facial expressions and speech.

[1214] Data transmission

[1215] The device transmits the captured video and audio data to a server via a network (e.g., Wi-Fi or LAN), thereby achieving real-time data collection and transmission.

[1216] Data analysis

[1217] The server preprocesses the received video and audio data using OpenCV for video preprocessing and Librosa for audio preprocessing to remove noise and normalize the data.

[1218] The server uses an AI model (e.g., TensorFlow or PyTorch) to identify objects from video data and detect abnormal sounds from audio data.

[1219] At the same time, the server utilizes an emotion engine (e.g., Azure Face API) to recognize the user's emotions, which includes facial expression analysis and voice analysis.

[1220] Generate and send a workaround

[1221] The server generates an optimal response based on the identified object and the user's emotions, for example, if a bear is identified, a loud cracking sound will be generated and increased if the user expresses fear.

[1222] The server sends the generated command for the solution to the terminal, which includes the solution to be executed and its parameters.

[1223] Implementing the solution

[1224] The device will then take action based on the received command, for example, playing a repellent sound from a speaker (e.g., JBL Charge 3) or emitting a strong light using an LED light.

[1225] Recheck and re-address

[1226] Even after implementing a countermeasure, the device continues to monitor the situation within the area and captures video and audio. This data is then sent back to the server, which uses an emotion engine to check for changes in the user's emotions. If further action is required as a result of the analysis, an appropriate countermeasure is regenerated and sent to the device.

[1227] Specific examples

[1228] For example, if a bear appears near a house, the system operates as follows:

[1229] The terminals (cameras and microphones installed around the home) capture video and audio of the bear and send the data to a server.

[1230] The server preprocesses the received video and audio, and uses an AI model to identify bears. In parallel, an emotion engine recognizes the user's emotions (e.g., fear or surprise).

[1231] The server generates a loud cracking sound that is effective against bears, and generates and sends to the terminal a command to adjust the volume according to the user's sense of fear.

[1232] The device plays a loud crackling sound from the speaker (JBL Charge 3) to scare off bears, then continues to monitor the area to see if any further action is needed.

[1233] In this way, the system of the present invention can detect the appearance of pests and vermin at an early stage, ensuring the safety and psychological security of users.

[1234] Prompt Sentence Examples

[1235] "Please explain how the system works if a bear is in the vicinity of a residential area."

[1236] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1237] Step 1: Gather information

[1238] The device uses a camera and microphone to capture video and audio within a set area in real time. The input is the situation within the area and the user's facial expressions and comments, and the output is the captured video and audio data. Specifically, the camera detects movement within the area and collects video data, while the microphone collects audio within the area and generates audio data.

[1239] Step 2: Send data

[1240] The terminal sends the captured video and audio data to the server via the network. The input is the video and audio data obtained in step 1, and the output is the data sent to the server. Specifically, the terminal generates video and audio data packets and sends them to the server via the network. If an error occurs during data transmission, it attempts to resend them.

[1241] Step 3: Data analysis

[1242] The server analyzes the received video and audio data. First, as preprocessing, it uses OpenCV to remove noise and normalize the video data, and Librosa to remove noise and normalize the audio data. The input is the video and audio data sent in step 2, and the output is the preprocessed data. Specifically, it performs color adjustment and noise removal on the video data, and performs frequency analysis on the audio data.

[1243] Step 4: Identifying the Object and Emotion

[1244] The server uses the AI ​​model to identify objects from the preprocessed video data and uses an emotion engine to recognize the user's emotions from the audio and video data. The input is the data preprocessed in step 3, and the output is the identified objects and recognized emotion data. Specifically, the video data is input into the AI ​​model to identify the object (e.g., a bear), and the emotion engine analyzes the user's facial expressions and tone of voice to recognize their emotions.

[1245] Step 5: Generate solutions

[1246] The server generates an optimal response based on the identified object and the recognized user emotion. The input is the object and emotion data obtained in step 4, and the output is a command for the response. Specifically, it uses an AI algorithm to generate an effective response to the object (e.g., a loud crackling sound), and adjusts the parameters of the response (volume, duration, etc.) according to the user emotion.

[1247] Step 6: Submit your solution

[1248] The server sends the generated remedy command to the terminal. The input is the remedy command generated in step 5, and the output is the command sent to the terminal. Specifically, the remedy command is packetized and sent to the terminal via the network.

[1249] Step 7: Implementing the solution

[1250] The device executes countermeasures based on the received command. The input is the command sent from the server in step 6, and the output is the executed countermeasure. Specific actions include playing a specific repellent sound (e.g., a loud cracking sound) from the speaker and emitting a strong light from the LED light if necessary.

[1251] Step 8: Reassess and readdress

[1252] Even after the countermeasure is implemented, the device continues to monitor the situation within the area, capturing video and audio and sending it back to the server. The input is the new video and audio data after the countermeasure is implemented, and the output is the data sent back to the server. The server analyzes the data again and checks for changes in the user's emotions. The input is the data after the countermeasure is implemented, and the output is the results of the reanalysis and a new countermeasure, if necessary. If the analysis shows that further action is necessary, an appropriate countermeasure is regenerated and sent to the device.

[1253] (Application example 2)

[1254] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1255] In recent years, the importance of security in living and social environments has increased. However, existing security systems have limitations in their ability to detect intrusions by suspicious individuals and prevent the intrusion of pests and vermin. Furthermore, they are unable to respond flexibly to user emotional states, making it difficult to provide a sense of psychological security. Therefore, there is a need for a system that can replace conventional systems, monitor situations in real time, automatically generate appropriate countermeasures, and provide optimal countermeasures that take the user's emotional state into account.

[1256] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1257] In this invention, the server includes means for collecting video and audio using a camera and microphone, means for receiving the collected video and audio data, means for analyzing the received data and identifying the target, and means for recognizing the user's emotion and optimizing the content and strength of the countermeasure based on the emotion, thereby enabling more effective and flexible security measures.

[1258] A "camera" is a device for collecting video data.

[1259] A "microphone" is a device for collecting audio data.

[1260] The "collection means" is a means for acquiring video and audio using a camera and microphone.

[1261] The "receiving means" is a means for transmitting collected video and audio data to a server via a network.

[1262] The "analysis means" is a means used to analyze the received video and audio data and identify objects or abnormalities.

[1263] An "object" is a specific entity within an area, such as an obstacle, suspicious person, pest, or vermin identified by the analysis means.

[1264] The "automatic countermeasure generation means" is a means for automatically generating an optimal countermeasure for a specified object.

[1265] The "command transmission means" is a means for transmitting an instruction to execute the generated countermeasure to the terminal.

[1266] The "countermeasure execution means" is a means for executing a specific countermeasure based on a received command.

[1267] The "re-monitoring means" is a means for re-monitoring the situation after the countermeasure has been implemented and collecting video and audio data again.

[1268] "Emotion recognition means" refers to a means for analyzing and recognizing a user's emotions.

[1269] The "optimization method" is a method for optimizing the content and intensity of countermeasures based on the recognized user emotions.

[1270] This invention is a system that uses cameras and microphones to collect video and audio in areas where pests and vermin appear, analyzes the video and audio to identify obstacles and suspicious individuals, and automatically generates countermeasures. This system is combined with an emotion recognition engine that recognizes the user's emotions and optimizes the content and strength of the countermeasures.

[1271] System Program

[1272] The system consists of the following components:

[1273] 1. Terminal

[1274] It includes a camera for collecting video data and a microphone for collecting audio data, and the cameras and microphones are installed inside and outside the home.

[1275] 2. Server

[1276] Data receiving means: Receives video and audio data transmitted from the terminal.

[1277] Data analysis method: Analyzes received video and audio data to identify objects (vermin, vermin, intruders, etc.). Face detection and abnormal sound detection models using OpenCV are used for the analysis.

[1278] Emotion recognition: Recognizes the user's emotions from the collected data and optimizes the strength and content of countermeasures based on that data. For emotion recognition, an emotion recognition model using Keras is used.

[1279] Countermeasure generation means: Automatically generates the optimal countermeasure based on the analysis results. For example, if an intruder is detected, a command to play an acoustic alarm is generated.

[1280] Command sending means: Sends the generated instructions for the solution to the terminal.

[1281] 3. Users

[1282] The system collects the user's emotional data (e.g., fear, surprise, etc.) and takes appropriate action based on that emotion.

[1283] Hardware and software used

[1284] Camera: Hardware for collecting video footage.

[1285] Microphone: Hardware for collecting sound.

[1286] Server: A central processing unit for data analysis and emotion recognition. A server with Python, OpenCV, and Keras installed is used.

[1287] Acoustic alarm device: Used as one of the countermeasures.

[1288] Specific examples of processing

[1289] For example, if an intruder appears near a home, the system will:

[1290] The devices (cameras and microphones installed around the house) capture video and audio of intruders and send the data to a server.

[1291] The server preprocesses the received video and audio data and uses an AI model to identify intruders. In parallel, an emotion recognition engine recognizes the user's emotions (e.g., fear or surprise).

[1292] The server generates an effective acoustic alarm against an intruder, and generates and sends to the terminal a command to adjust the intensity of the alarm according to the user's emotion data.

[1293] The device will play an audible alarm to scare off intruders, then continue monitoring to see if any further action is required.

[1294] Examples of prompt statements

[1295] "Design a smartphone application that runs inside the home. It uses a camera and microphone to capture video and audio in real time to detect intruders and suspicious activity. It uses an emotion recognition engine to detect user emotions (especially fear and surprise), and if an abnormality is detected, it will sound an alarm and notify the user or security company. Using Python, OpenCV, and Keras, you will pre-train the emotion recognition model and perform real-time analysis."

[1296] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1297] Step 1:

[1298] The device uses a camera and microphone to collect video and audio data in real time. This data includes the situation within the area, the user's facial expressions, and their comments. The collected data serves as basic information for understanding the situation within the area.

[1299] Step 2:

[1300] The terminals transmit the collected video and audio data to a server via a network. This transmission is performed in real time, and it is important to minimize delays.

[1301] Step 3:

[1302] The server preprocesses the received video and audio data, specifically removing noise and normalizing the data. This preprocessing improves data quality and increases analysis accuracy.

[1303] Step 4:

[1304] The server analyzes the preprocessed data. It uses an AI model to identify objects (such as suspicious people or pests) from the video data, and detects abnormal sounds (such as the sound of destruction or metal) from the audio data. The analysis results are used to confirm the presence of the object.

[1305] Step 5:

[1306] The server uses an emotion recognition engine to analyze the user's emotions. By analyzing facial expressions from video data and changes in tone and volume from audio data, it can determine what emotions the user is feeling (for example, fear or surprise).

[1307] Step 6:

[1308] The server generates the optimal countermeasure based on the analysis results. Specifically, it automatically determines the best action to take against the target object (for example, playing a loud cracking sound or emitting ultrasound). It also adjusts the strength and content of the countermeasure according to the user's emotions.

[1309] Step 7:

[1310] The server sends the generated instructions (commands) for the countermeasures to the terminal. These commands include the countermeasures to be executed and their parameters (e.g., type and intensity of sound, light pattern, etc.).

[1311] Step 8:

[1312] The device will then take action based on the command it receives. For example, if it receives a voice command, it will play a repellent sound from a specific speaker, and if it receives a video command, it will emit a strong light. This will allow it to take action against the target.

[1313] Step 9:

[1314] Even after the countermeasure is implemented, the device continues to monitor the situation within the area, collecting video and audio data again and sending it to the server.

[1315] Step 10:

[1316] The server analyzes the data again to see how the user's emotions have changed. Based on the analysis results, it determines whether further action is necessary. If the object still exists and further action is necessary, the process repeats from step 6.

[1317] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1318] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1319] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1320] [Fourth embodiment]

[1321] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1322] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1323] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1324] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1325] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1326] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1327] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1328] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1329] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1330] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1331] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1332] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1333] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1334] ---

[1335] MODE FOR CARRYING OUT THE INVENTION

[1336] The present invention relates to a system for detecting the appearance of pests and vermin at an early stage and automatically taking appropriate measures. This system includes a series of processes that use a camera and a microphone to collect video and audio within an area, analyze the video and audio to identify the target, and take appropriate measures. The following describes in detail an embodiment of the present invention.

[1337] Overall system overview

[1338] This system consists of cameras and microphones (hereafter referred to as terminals) installed within the area, and a server that receives and analyzes the data sent from these terminals and implements countermeasures.

[1339] Program processing

[1340] 1. Information gathering

[1341] The device captures video and audio within the set area in real time, and this data is used as the basis for understanding the situation within the area.

[1342] 2. Data Transmission

[1343] The device transmits the captured video and audio data to a server via a network, enabling real-time data collection and transmission.

[1344] 3. Data Analysis

[1345] The server analyzes the received video and audio data, and performs preprocessing such as noise removal and data normalization.

[1346] The server uses an AI model (e.g., an image recognition model or a voice recognition model using deep learning) to identify the object (e.g., a bear or a hornet).

[1347] 4. Generating solutions

[1348] The server automatically generates the most effective response to the object identified through the analysis: for example, a loud cracking sound for a bear, or ultrasonic waves for a hornet.

[1349] 5. Sending solutions

[1350] The server sends the generated command for the solution to the terminal, which includes the specific solution and the timing for its execution.

[1351] 6. Implementing the solution

[1352] The device will then take action based on the command it receives, for example, playing a specific sound or emitting a bright light.

[1353] 7. Recheck and Readdress

[1354] After the countermeasure is implemented, the device continues to monitor the situation within the area and captures video and audio again, which are then sent to the server for further analysis.

[1355] If the server determines that further action is necessary, it generates an appropriate countermeasure and sends it to the device. By repeating this process, the target object is completely avoided.

[1356] Specific examples

[1357] For example, if a bear appears near a house, the system operates as follows:

[1358] The terminals (cameras installed around the home) capture images of bears and send the data to a server.

[1359] The server preprocesses the received footage and uses an AI model to identify the bear.

[1360] The server generates a command to play a loud cracking sound that is effective against bears and sends it to the device.

[1361] The device plays a loud cracking sound to scare off bears.

[1362] The device will capture the video again to check if the bear has completely left the area, and take further action if necessary.

[1363] Similarly, if a hornet's nest is discovered on the eaves of a house, the device captures video and audio, and the server analyzes them to identify the nest. The server selects a countermeasure to generate ultrasonic waves that repel hornets and sends it to the device. The device repels hornets by emitting ultrasonic waves.

[1364] This system allows users to detect the appearance of pests and vermin early and deal with them effectively, thereby providing a safe living environment. This series of processes is fully automated, so users do not need any special technical knowledge or operation to operate the system.

[1365] The above is a specific embodiment for carrying out the present invention.

[1366] The processing flow will be explained below.

[1367] Step-by-step process

[1368] Step 1:

[1369] The device captures video and audio within the set area in real time. Video is captured using a camera and audio is collected using a microphone. The collected data is temporarily stored in local storage and updated periodically.

[1370] Step 2:

[1371] The devices transmit the collected video and audio data over a network to a server. This data is streamed in real time or sent in batches. The data transmission is encrypted to ensure security.

[1372] Step 3:

[1373] The server receives the video and audio data sent from the device, saves the data in the data storage, and copies it to the workspace for analysis.

[1374] Step 4:

[1375] The server performs preprocessing on the received data, specifically noise removal, data normalization, and missing value imputation, to improve data quality and enhance analysis accuracy.

[1376] Step 5:

[1377] The server inputs the preprocessed data into the AI ​​model to detect objects. It uses an image recognition model to identify objects from video data and a voice recognition model to detect abnormal sounds from audio data.

[1378] Step 6:

[1379] The server automatically generates the most appropriate response to the object identified through the analysis: for example, a loud cracking sound is generated if a bear is identified, and ultrasonic waves are generated if a hornet is identified.

[1380] Step 7:

[1381] The server sends the generated command for the countermeasure to the terminal. This command includes the countermeasure to be executed and its parameters (e.g., type and intensity of sound, light pattern, etc.).

[1382] Step 8:

[1383] The device will then take action based on the command it receives, for example, playing a repellent sound from a specific speaker if a voice command is received, or emitting a bright light if a video command is received.

[1384] Step 9:

[1385] The device will monitor the situation again after the countermeasure is implemented, capturing video and audio data, which will also be sent to the server again.

[1386] Step 10:

[1387] The server analyzes the retransmitted data and determines whether additional action is required. If the object is still present, it generates an appropriate action again and sends it to the device. Repeat from step 6 if necessary.

[1388] The above is the flow of the program processing including the specific operations at each step.

[1389] Example 1

[1390] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1391] Conventional systems for detecting and responding to the emergence of pests and vermin have problems with real-time accuracy, often resulting in delayed responses and insufficient effectiveness. Furthermore, manual operation is required, and users without special technical knowledge have had the problem of being unable to operate the systems properly. Furthermore, the methods offered as avoidance measures by the systems are uniform, making it difficult to provide optimal countermeasures tailored to the characteristics of the target.

[1392] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1393] In this invention, the server includes a means for analyzing collected video and audio data to identify an object, a means for automatically generating a countermeasure corresponding to the identified object, and a means for transmitting a command for executing the generated countermeasure, thereby enabling high-precision detection of an object in real time and automatic appropriate countermeasures.

[1394] An "image capture device" is a device installed to capture video within a specific area in real time.

[1395] "Audio recording device" refers to equipment installed to record audio within a specific area.

[1396] "Means for collecting" refers to the function or process of acquiring video and audio data using image capture devices and audio recording devices.

[1397] The "means for receiving" is a function or process for transmitting collected video and audio data to a server via a network and for the server to receive it.

[1398] "Means for analyzing" refers to algorithms or systems that process received video and audio data and identify objects.

[1399] The "means for automatic generation" is a function or process that automatically creates an appropriate countermeasure for the target object based on the results of data analysis.

[1400] The "means for sending a command" is a function or process that sends an instruction to execute the generated countermeasure to the terminal via the network.

[1401] The "means for executing" refers to a function or device that the terminal actually physically executes the countermeasure based on the received command.

[1402] "Means for monitoring and analysis" refers to the function or process of rechecking the situation after the countermeasures have been implemented and reanalyzing it based on newly collected data.

[1403] The "pre-processing means" refers to a function or process that performs noise removal and data normalization on the collected video and audio data.

[1404] A "generative AI model" is a model or algorithm that uses deep learning or machine learning to identify objects from data.

[1405] MODE FOR CARRYING OUT THE INVENTION

[1406] The present invention provides a system for detecting the appearance of pests and vermin at an early stage and automatically taking appropriate countermeasures. This system includes a series of processes that use an image capture device and an audio recording device to collect video and audio within an area, analyze the video and audio, identify the target, and take appropriate countermeasures. The following describes in detail the embodiments of the present invention.

[1407] Overall system overview

[1408] This system consists of image capture devices and audio recording devices (hereafter referred to as terminals) installed within the area, and a server that receives and analyzes the data sent from these terminals and executes countermeasures. The system is designed to be easy to operate, even if the user does not have special technical knowledge.

[1409] Program processing

[1410] 1. Information collection via devices

[1411] The device captures video and audio within the set area in real time. The image capture device captures video at 30 frames per second, and the audio recorder records audio at 24-bit / 48kHz.

[1412] 2. Data transmission by the terminal

[1413] The device periodically compresses the captured video and audio data, converts it into a data structure such as JSON or Protobuf format, and sends it to the server using HTTP / HTTPS or WebSocket protocols.

[1414] 3. Data analysis by the server (preprocessing)

[1415] The server first decodes the received data and performs noise reduction on the video and audio. The video data is decoded using the OpenCV library to reduce blur and noise, and the audio data is decoded using Python's pydub and SciPy.

[1416] 4. Data analysis by the server (identification of the target object)

[1417] The server analyzes the video and audio data using a generative AI model (for example, a deep learning model using TensorFlow or PyTorch) and uses object detection algorithms such as YOLO (You Only Look Once) or R-CNN to identify objects such as bears or hornets.

[1418] 5. Server-generated solutions

[1419] Based on the identified object, the server uses a pre-defined rules-based system to generate the optimal response, such as playing a loud cracking sound for a bear or generating ultrasound for a hornet.

[1420] 6. Server sends solution

[1421] The server sends the generated command for the solution to the appropriate device at the appropriate time. This command includes the specific content of the solution (e.g., the path of the audio file, playback timing, etc.).

[1422] 7. Implementing measures on the device

[1423] The device will then take action based on the command it receives, such as playing a loud crackling sound through the speaker or flashing an LED light.

[1424] 8. Reconfirmation and re-action by terminal

[1425] After implementing a countermeasure, the device continues to monitor the situation within the area and recaptures video and audio data. This data is then sent to the server for further analysis. If the server determines that further countermeasures are necessary, it generates a new appropriate countermeasure and sends it to the device. This process is repeated as many times as necessary.

[1426] Specific examples

[1427] For example, if a bear appears near a house, the system operates as follows:

[1428] 1. Information collection via devices

[1429] The device captures video at 30 frames per second using an image capture device installed in the garden, and records audio at 24-bit / 48kHz using an audio recording device.

[1430] 2. Data transmission by the terminal

[1431] The device compresses the captured data every 5 seconds, converts it into JSON format, and sends it to the server using the HTTP protocol.

[1432] 3. Data analysis by the server (preprocessing)

[1433] The server uses OpenCV to remove noise from the received video data and pydub to perform noise filtering on the audio data.

[1434] 4. Data analysis by the server (identification of the target object)

[1435] The server uses the YOLO model to analyze the video data and identify the presence of bears.

[1436] 5. Server-generated solutions

[1437] The server generates commands based on pre-defined rules to play a loud cracking sound that works on bears.

[1438] 6. Server sends solution

[1439] The server sends the generated command (e.g., the path to the audio file and the playback timing) to the terminal.

[1440] 7. Implementing measures on the device

[1441] Based on the command received, the device plays a loud cracking sound through its speaker to scare off bears.

[1442] 8. Reconfirmation and re-action by terminal

[1443] The device captures the video again to check if the bear has completely left the area. If the bear is still in the area, the server generates a new countermeasure and sends it to the device to execute again.

[1444] Example prompts for generative AI models

[1445] "Please tell me how to use this system to automatically detect the appearance of bears and take appropriate action."

[1446] This is a specific mode for carrying out the invention.

[1447] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1448] Program processing steps

[1449] Step 1: Gather information

[1450] The device captures video and audio within the area in real time. Specifically, the image capture device captures video at 30 frames per second, and the audio recording device records audio at 24-bit / 48kHz.

[1451] Input: Physical conditions within the set area

[1452] Output: Video and audio data

[1453] Step 2: Send data

[1454] The device periodically compresses the captured video and audio data, converts it into JSON or Protobuf format, and sends it to the server using HTTP / HTTPS or WebSocket protocol. For example, the device compresses and sends the data every 5 seconds.

[1455] Input: Video and audio data

[1456] Output: Compressed data

[1457] Step 3: Data analysis (preprocessing)

[1458] The server decodes the received data and performs noise reduction on the video and audio. Specifically, the OpenCV library is used to reduce blur and noise in the video data, and pydub and SciPy are used to perform noise filtering on the audio data.

[1459] Input: Compressed data

[1460] Output: Denoised data

[1461] Step 4: Data analysis (target identification)

[1462] The server analyzes the video and audio data using a deep learning model (e.g., a generative AI model using TensorFlow or PyTorch) to identify objects using object detection algorithms such as YOLO or R-CNN.

[1463] Input: Denoised data

[1464] Output: Information about the identified object (e.g., location and type of bear or hornet)

[1465] Step 5: Generate solutions

[1466] Based on the identified object, the server uses a rules engine to automatically generate the optimal response, for example, a loud cracking sound for a bear or ultrasonic waves for a hornet.

[1467] Input: Information about the identified object

[1468] Output: Action command

[1469] Step 6: Submit your solution

[1470] The server generates a command for the appropriate action and sends it to the device at the appropriate time, which includes specific instructions such as the path to the audio file and the timing of playback.

[1471] Input: Solution command

[1472] Output: Commands sent to the terminal

[1473] Step 7: Implementing the solution

[1474] The device will then take action based on the command it receives, such as playing a loud crackling sound through the speaker or flashing an LED light.

[1475] Input: Command sent to the terminal

[1476] Output: Actions taken

[1477] Step 8: Reassess and readdress

[1478] Even after the device has implemented the countermeasures, it continues to monitor the situation within the area, recapturing video and audio data and sending it to the server. If it determines that reanalysis is necessary, it generates an appropriate countermeasure again and sends it to the device.

[1479] Input: Post-execution video and audio data

[1480] Output: Updated workaround command

[1481] The above is the processing flow of the specific program of the system.

[1482] (Application example 1)

[1483] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1484] Automating the early detection of pests and effective countermeasures in areas where pests and vermin appear is a major challenge in terms of safety management of living environments and workplaces. Conventional systems often detect and counter these pests manually, making it difficult to respond quickly. Furthermore, there is a lack of technology to quickly implement optimal countermeasures for specific pests.

[1485] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1486] In this invention, the server includes means for collecting video and audio using an imaging device and an audio collecting device installed in an area where pests or vermin appear, means for transmitting the collected video and audio data to a mobile communication terminal, means for analyzing the data received by the mobile communication terminal and identifying targets, means for automatically generating countermeasures corresponding to the identified targets, means for transmitting instructions for executing the generated countermeasures, means for executing the countermeasures based on the received instructions, and means for re-monitoring and analyzing the situation after execution. This enables early detection of pests and rapid and effective countermeasures.

[1487] - "Imaging device" refers to hardware for capturing images.

[1488] "Sound collection device" refers to hardware for collecting sound.

[1489] "Mobile communication terminal" refers to a terminal that uses a mobile communication network, such as a mobile phone or smartphone.

[1490] "Analyzing the data" refers to the process of processing collected video and audio data to identify specific information.

[1491] "Target" refers to pests or vermin identified from the analyzed data.

[1492] "Automatically generating countermeasures" refers to the process by which the system automatically determines the optimal countermeasure based on the identified objectives.

[1493] "Directives" refer to the specific commands required to execute the generated countermeasures.

[1494] "Implementing countermeasures" refers to the process of actually taking countermeasures based on instructions.

[1495] "Re-monitoring and analyzing" refers to the process of collecting the situation after the countermeasures have been implemented and performing the same analysis.

[1496] The present invention relates to a system for early detection and rapid response in areas where pests or vermin appear. This system includes a series of processes that collect video and audio using an imaging device and an audio collecting device, transmit the video and audio to a mobile communication terminal for analysis, identify targets, and automatically generate and execute countermeasures.

[1497] Overall system overview

[1498] The system consists of the following major components:

[1499] 1. Imaging and audio collection devices: These are devices that are installed in the target area and continuously collect video and audio.

[1500] 2. Mobile communication terminal: Usually a smartphone, which receives the collected data and analyzes it using AI models.

[1501] 3. Server: Communicates with mobile communication terminals and sends commands.

[1502] Program processing procedure

[1503] The server first receives the video and audio transmitted from the imaging and audio collection devices. It then preprocesses the received data, removing noise and normalizing it. This process is performed using software such as OpenCV (for video processing) and Librosa (for audio processing).

[1504] The server then uses AI models to perform image and speech recognition, using deep learning libraries such as Tensorflow and Keras. For example, YOLO and Faster R-CNN are used for image recognition, and WaveNet and Transformer-based models are used for speech recognition.

[1505] Once a target is identified, the server automatically generates a countermeasure accordingly: for example, if a bear is identified, a countermeasure that plays a loud cracking sound is generated, and if a hornet is identified, a countermeasure that generates ultrasound is selected.

[1506] The generated countermeasures are sent as commands to the mobile communication device, which then executes the actual countermeasures based on the received commands, such as playing a sound from the smartphone speaker or emitting light or ultrasound from a specific device.

[1507] Specific examples

[1508] For example, if a bear appears near a house, the system operates as follows:

[1509] The imaging and sound collection devices capture images and sounds of the bear and transmit the data to a mobile communication terminal.

[1510] The mobile communication terminal receives the data and performs noise removal and normalization processing.

[1511] Next, an analysis is performed using an AI model to identify the bear.

[1512] Once a bear is identified, a command is sent from the server to the mobile communication terminal to play a loud cracking sound.

[1513] The mobile device will then play a clicking sound to scare off the bears, and will continue to monitor the area and take further action if necessary.

[1514] This system allows users to automatically detect pests and vermin early and take swift and effective measures to maintain a safe living and working environment.

[1515] Prompt Sentence Examples

[1516] "Please create a program that analyzes camera footage and audio from around the home to detect bears and hornets. Please also include a function that automatically suggests and executes appropriate countermeasures when specific pests or vermin are detected."

[1517] The above is an embodiment of the present invention.

[1518] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1519] Step 1: Gather information

[1520] The terminal uses an imaging device and an audio collecting device to capture video and audio within the area in real time.

[1521] Input: Raw video and audio data captured from imaging and audio collection devices.

[1522] Data processing: Format conversion and temporary storage of captured data.

[1523] Output: Format-converted and temporarily stored video and audio data.

[1524] Step 2: Send data

[1525] The terminal transmits the captured video and audio data to a server via a network.

[1526] Input: Format-converted and buffered video and audio data.

[1527] Data processing: Packetizing data and sending it over the network.

[1528] Output: Data packets sent to the server.

[1529] Step 3: Data reception and preprocessing

[1530] The server receives the data packets sent from the terminal and restores them as video and audio data, as well as performs noise removal and data normalization.

[1531] Input: Data packets sent from the terminal.

[1532] Data processing: Reformatting data packets, removing noise, and normalizing data.

[1533] Output: Pre-processed video and audio data.

[1534] Step 4: Data analysis and target identification

[1535] The server analyzes the pre-processed video and audio data using a generative AI model to identify targets, using deep learning models such as YOLO, Faster R-CNN, and WaveNet using Tensorflow and Keras.

[1536] Input: Preprocessed video and audio data.

[1537] Data computation: Deep learning analysis using AI models.

[1538] Output: Information about identified targets (bears, hornets, etc.).

[1539] Step 5: Generate solutions

[1540] The server automatically generates the best response depending on the identified target, for example, a loud cracking sound for a bear or ultrasonic waves for a hornet.

[1541] Input: Identified landmark information.

[1542] Data calculation: Execution of algorithm for automatic generation of countermeasures.

[1543] Output: Generated action instructions.

[1544] Step 6: Send instructions for how to deal with the problem

[1545] The server sends a command to the terminal to execute the generated countermeasure.

[1546] Input: Generated action instructions.

[1547] Data processing: Packetizing commands and sending them over the network.

[1548] Output: Command packets sent to the terminal.

[1549] Step 7: Implementing the solution

[1550] The device then takes action based on the received command, for example by playing a specific sound from a smartphone speaker or emitting light or ultrasound from a specific device.

[1551] Input: Received command packet.

[1552] Specific actions: playing sound, emitting light, generating ultrasound.

[1553] Output: The results of the action taken.

[1554] Step 8: Remonitor and reanalyze

[1555] After implementing the countermeasures, the device continues to monitor the area and transmits newly captured video and audio data to the server again. The server analyzes the data again and, if necessary, implements the countermeasures again.

[1556] Input: Newly captured video and audio data.

[1557] Data calculation: Pre-processing for reanalysis and analysis using AI models.

[1558] Specific actions: Continuous monitoring and repeated analysis within the area.

[1559] Output: Decision whether the situation is safe or whether further action is required.

[1560] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1561] ---

[1562] MODE FOR CARRYING OUT THE INVENTION

[1563] The present invention combines a system that uses a camera and a microphone to collect video and audio within an area, analyzes the video and audio to identify the target, and automatically executes an appropriate countermeasure, with an emotion engine that recognizes the user's emotions. This emotion engine makes it possible to optimize and improve the countermeasure, resulting in a more effective response. The following describes in detail the embodiments of the present invention.

[1564] Overall system overview

[1565] This system consists of cameras and microphones (hereinafter referred to as terminals) installed within the area, a server that receives and analyzes the data sent from these terminals and implements countermeasures, and an emotion engine that recognizes the user's emotions.

[1566] Program processing

[1567] 1. Information gathering

[1568] The device captures video and audio within the designated area in real time. This data is used as basic information to understand the situation within the area. The device also captures the user's facial expressions and comments at the same time.

[1569] 2. Data Transmission

[1570] The device transmits the captured video and audio data to a server via a network, enabling real-time data collection and transmission.

[1571] 3. Data Analysis

[1572] The server analyzes the received video and audio data, and performs preprocessing such as noise removal and data normalization.

[1573] The server uses an AI model to identify objects from video data and detect abnormal sounds from audio data.

[1574] At the same time, the server utilizes an emotion engine to recognize the user's emotions, which includes facial expression analysis and voice analysis.

[1575] 4. Generating solutions

[1576] The server automatically generates the most appropriate response to the object identified through the analysis: for example, a loud cracking sound is generated if a bear is identified, and ultrasonic waves are generated if a hornet is identified.

[1577] The content and strength of the coping strategies are optimized based on the user's recognized emotions. For example, if the user expresses fear, a stronger coping strategy is selected.

[1578] 5. Sending solutions

[1579] The server sends the generated command for the countermeasure to the terminal. This command includes the countermeasure to be executed and its parameters (e.g., type and intensity of sound, light pattern, etc.).

[1580] 6. Implementing the solution

[1581] The device will then take action based on the command it receives, for example, playing a repellent sound from a specific speaker if it receives a voice command, or emitting a bright light if it receives a video command.

[1582] 7. Recheck and Readdress

[1583] After the countermeasure is implemented, the device continues to monitor the situation within the area and captures video and audio data, which is also sent to the server again.

[1584] The server uses the emotion engine to analyze the transmitted data again and check how the user's emotions have changed.

[1585] Based on the analysis, it is determined whether additional measures are necessary. If the object is still present, appropriate measures are generated again and sent to the device. Repeat from step 4 as necessary.

[1586] Specific examples

[1587] For example, if a bear appears near a house, the system operates as follows:

[1588] The terminal (a camera installed around the house) captures video of the bear and sends that data along with the user's video and audio data to the server.

[1589] The server preprocesses the received video and uses an AI model to identify bears, while an emotion engine recognizes the user's emotions (e.g., fear or surprise).

[1590] The server generates a loud cracking sound that is effective against bears, and generates and sends to the terminal a command to adjust the volume according to the user's sense of fear.

[1591] The device will play a loud crackling sound to scare off the bears, and will continue to monitor the situation to see if any further action is needed.

[1592] Similarly, if a hornet's nest is found on the eaves of a house, the device captures video and audio, and the server analyzes them to identify the nest. The server generates ultrasonic waves that repel hornets and sends a command to the device to emit them at an intensity that corresponds to the user's emotions (e.g., anxiety). The device repels hornets by emitting ultrasonic waves.

[1593] This system allows users to detect the appearance of pests and vermin early and deal with them effectively, thereby providing a safe living environment. It also provides a sense of psychological security by optimizing countermeasures that take into account the user's emotions. This entire process is fully automated, so users do not need any special technical knowledge or operations to operate the system.

[1594] The above is a specific embodiment for carrying out the present invention.

[1595] The processing flow will be explained below.

[1596] Step-by-step process

[1597] Step 1:

[1598] The device captures video and audio within the set area in real time. Video is captured using a camera, and audio is collected using a microphone. The device also captures the user's facial expressions and voice. This data is temporarily stored in local storage and updated periodically.

[1599] Step 2:

[1600] The devices transmit the collected video and audio data over a network to a server. This data is streamed in real time or sent in batches. The data transmission is encrypted to ensure security.

[1601] Step 3:

[1602] The server receives the video and audio data sent from the device, saves the data in the data storage, and copies it to the workspace for analysis.

[1603] Step 4:

[1604] The server performs preprocessing on the received data, specifically noise removal, data normalization, and missing value imputation, to improve data quality and enhance analysis accuracy.

[1605] Step 5:

[1606] The server inputs the preprocessed data into the AI ​​model to detect objects. It uses an image recognition model to identify objects from video data and a voice recognition model to detect abnormal sounds from audio data.

[1607] Step 6:

[1608] The server uses an emotion engine to analyze the user's facial expressions and voice to recognize the user's emotions, such as surprise, fear, relief, etc.

[1609] Step 7:

[1610] The server automatically generates the most appropriate response to the object identified through the analysis: for example, a loud cracking sound is generated if a bear is identified, and ultrasonic waves are generated if a hornet is identified.

[1611] The content and strength of the coping strategies are optimized based on the user's recognized emotions. For example, if the user expresses fear, the system will adjust the frequency of the coping strategies.

[1612] Step 8:

[1613] The server sends the generated command for the countermeasure to the terminal. This command includes the countermeasure to be executed and its parameters (e.g., type and intensity of sound, light pattern, etc.).

[1614] Step 9:

[1615] The device will then take action based on the command it receives, for example, playing a repellent sound from a specific speaker if a voice command is received, or emitting a bright light if a video command is received.

[1616] Step 10:

[1617] The device will monitor the situation again after the countermeasure is implemented, capturing video and audio data, which will also be sent to the server again.

[1618] Step 11:

[1619] The server analyzes the retransmitted data and determines whether additional measures are necessary. If changes are observed in the object and the user's emotions have improved, the server determines that the measures are complete.

[1620] The server uses the emotion engine to analyze the transmitted data again and check how the user's emotions have changed.

[1621] If the server determines that further action is required, it generates an appropriate solution again and sends it to the terminal.

[1622] The above is the flow of program processing including specific operations.

[1623] Example 2

[1624] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1625] Conventional pest and vermin control systems have the problem of being unable to consider the psychological state of users, who have emotions, making it difficult to select the optimal countermeasure. Furthermore, conventional systems have low accuracy in identifying targets, making them prone to misidentification or failure to recognize them. This not only fails to adequately ensure user safety, but can also result in unnecessary countermeasures being taken. Furthermore, they lack the functionality to provide appropriate feedback on changes in the situation or the user's emotions after implementing countermeasures.

[1626] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recognizing the user's emotions, a means for automatically generating a countermeasure according to the identified object and the user's emotions, and a means for re-monitoring and analyzing the situation after execution and the user's emotions. This allows for the selection of an appropriate and effective countermeasure for the object, and also enables an optimal response that takes the user's emotions into consideration. As a result, the user's safety and psychological sense of security can be improved.

[1627] "Pests and vermin" refers to animals or insects that may appear in an area and cause human or physical harm.

[1628] "Camera" refers to a photographic device for capturing video data.

[1629] "Microphone" refers to an acoustic receiving device for capturing audio data.

[1630] "Collection means" refers to a combination of equipment and software for collecting video and audio data from within a specified area.

[1631] "Means for receiving" refers to the communications network and software for transmitting collected video and audio data to the server.

[1632] "Means for analyzing and identifying objects" refers to the processes and devices that use AI models and algorithms to identify specific pests or vermin from the data received.

[1633] "Means for recognizing user emotions" refers to a combination of software and hardware for analyzing the user's facial expressions and voice and determining their psychological state.

[1634] "Means for automatically generating countermeasures" refers to algorithms and programs for deriving optimal countermeasures based on the identified object and the user's emotions.

[1635] "Means for sending commands" refers to a communication network and software that sends instructions to the terminal to execute the generated countermeasure.

[1636] "Means for implementing countermeasures" refers to the equipment and software that implements the actual countermeasures (such as sound or light patterns) within the area based on the transmitted commands.

[1637] "Means for re-monitoring and re-analysis" refers to a combination of equipment and software for continuously monitoring the situation in an area after countermeasures have been implemented and for re-analyzing as necessary.

[1638] The present invention is a system that uses a camera and a microphone to collect and analyze video and audio within an area, identify the object, and automatically execute appropriate countermeasures. This system incorporates an emotion engine that recognizes the user's emotions, enabling optimization and improvement of countermeasures to achieve more effective responses. Detailed embodiments of the present invention are described below.

[1639] System Overview

[1640] This system consists of cameras and microphones (hereinafter referred to as terminals) installed within the area, a server that receives and analyzes the data sent from these terminals and implements countermeasures, and an emotion engine that recognizes the user's emotions.

[1641] Information gathering

[1642] The device captures video and audio within a set area in real time using a camera (e.g., Logitech C920) and a microphone (e.g., Blue Yeti). The collected data includes the objects within the area as well as the user's facial expressions and speech.

[1643] Data transmission

[1644] The device transmits the captured video and audio data to a server via a network (e.g., Wi-Fi or LAN), thereby achieving real-time data collection and transmission.

[1645] Data analysis

[1646] The server preprocesses the received video and audio data using OpenCV for video preprocessing and Librosa for audio preprocessing to remove noise and normalize the data.

[1647] The server uses an AI model (e.g., TensorFlow or PyTorch) to identify objects from video data and detect abnormal sounds from audio data.

[1648] At the same time, the server utilizes an emotion engine (e.g., Azure Face API) to recognize the user's emotions, which includes facial expression analysis and voice analysis.

[1649] Generate and send a workaround

[1650] The server generates an optimal response based on the identified object and the user's emotions, for example, if a bear is identified, a loud cracking sound will be generated and increased if the user expresses fear.

[1651] The server sends the generated command for the solution to the terminal, which includes the solution to be executed and its parameters.

[1652] Implementing the solution

[1653] The device will then take action based on the received command, for example, playing a repellent sound from a speaker (e.g., JBL Charge 3) or emitting a strong light using an LED light.

[1654] Recheck and re-address

[1655] Even after implementing a countermeasure, the device continues to monitor the situation within the area and captures video and audio. This data is then sent back to the server, which uses an emotion engine to check for changes in the user's emotions. If further action is required as a result of the analysis, an appropriate countermeasure is regenerated and sent to the device.

[1656] Specific examples

[1657] For example, if a bear appears near a house, the system operates as follows:

[1658] The terminals (cameras and microphones installed around the home) capture video and audio of the bear and send the data to a server.

[1659] The server preprocesses the received video and audio, and uses an AI model to identify bears. In parallel, an emotion engine recognizes the user's emotions (e.g., fear or surprise).

[1660] The server generates a loud cracking sound that is effective against bears, and generates and sends to the terminal a command to adjust the volume according to the user's sense of fear.

[1661] The device plays a loud crackling sound from the speaker (JBL Charge 3) to scare off bears, then continues to monitor the area to see if any further action is needed.

[1662] In this way, the system of the present invention can detect the appearance of pests and vermin at an early stage, ensuring the safety and psychological security of users.

[1663] Prompt Sentence Examples

[1664] "Please explain how the system works if a bear is in the vicinity of a residential area."

[1665] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1666] Step 1: Gather information

[1667] The device uses a camera and microphone to capture video and audio within a set area in real time. The input is the situation within the area and the user's facial expressions and comments, and the output is the captured video and audio data. Specifically, the camera detects movement within the area and collects video data, while the microphone collects audio within the area and generates audio data.

[1668] Step 2: Send data

[1669] The terminal sends the captured video and audio data to the server via the network. The input is the video and audio data obtained in step 1, and the output is the data sent to the server. Specifically, the terminal generates video and audio data packets and sends them to the server via the network. If an error occurs during data transmission, it attempts to resend them.

[1670] Step 3: Data analysis

[1671] The server analyzes the received video and audio data. First, as preprocessing, it uses OpenCV to remove noise and normalize the video data, and Librosa to remove noise and normalize the audio data. The input is the video and audio data sent in step 2, and the output is the preprocessed data. Specifically, it performs color adjustment and noise removal on the video data, and performs frequency analysis on the audio data.

[1672] Step 4: Identifying the Object and Emotion

[1673] The server uses the AI ​​model to identify objects from the preprocessed video data and uses an emotion engine to recognize the user's emotions from the audio and video data. The input is the data preprocessed in step 3, and the output is the identified objects and recognized emotion data. Specifically, the video data is input into the AI ​​model to identify the object (e.g., a bear), and the emotion engine analyzes the user's facial expressions and tone of voice to recognize their emotions.

[1674] Step 5: Generate solutions

[1675] The server generates an optimal response based on the identified object and the recognized user emotion. The input is the object and emotion data obtained in step 4, and the output is a command for the response. Specifically, it uses an AI algorithm to generate an effective response to the object (e.g., a loud crackling sound), and adjusts the parameters of the response (volume, duration, etc.) according to the user emotion.

[1676] Step 6: Submit your solution

[1677] The server sends the generated remedy command to the terminal. The input is the remedy command generated in step 5, and the output is the command sent to the terminal. Specifically, the remedy command is packetized and sent to the terminal via the network.

[1678] Step 7: Implementing the solution

[1679] The device executes countermeasures based on the received command. The input is the command sent from the server in step 6, and the output is the executed countermeasure. Specific actions include playing a specific repellent sound (e.g., a loud cracking sound) from the speaker and emitting a strong light from the LED light if necessary.

[1680] Step 8: Reassess and readdress

[1681] Even after the countermeasure is implemented, the device continues to monitor the situation within the area, capturing video and audio and sending it back to the server. The input is the new video and audio data after the countermeasure is implemented, and the output is the data sent back to the server. The server analyzes the data again and checks for changes in the user's emotions. The input is the data after the countermeasure is implemented, and the output is the results of the reanalysis and a new countermeasure, if necessary. If the analysis shows that further action is necessary, an appropriate countermeasure is regenerated and sent to the device.

[1682] (Application example 2)

[1683] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1684] In recent years, the importance of security in living and social environments has increased. However, existing security systems have limitations in their ability to detect intrusions by suspicious individuals and prevent the intrusion of pests and vermin. Furthermore, they are unable to respond flexibly to user emotional states, making it difficult to provide a sense of psychological security. Therefore, there is a need for a system that can replace conventional systems, monitor situations in real time, automatically generate appropriate countermeasures, and provide optimal countermeasures that take the user's emotional state into account.

[1685] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1686] In this invention, the server includes means for collecting video and audio using a camera and microphone, means for receiving the collected video and audio data, means for analyzing the received data and identifying the target, and means for recognizing the user's emotion and optimizing the content and strength of the countermeasure based on the emotion, thereby enabling more effective and flexible security measures.

[1687] A "camera" is a device for collecting video data.

[1688] A "microphone" is a device for collecting audio data.

[1689] The "collection means" is a means for acquiring video and audio using a camera and microphone.

[1690] The "receiving means" is a means for transmitting collected video and audio data to a server via a network.

[1691] The "analysis means" is a means used to analyze the received video and audio data and identify objects or abnormalities.

[1692] An "object" is a specific entity within an area, such as an obstacle, suspicious person, pest, or vermin identified by the analysis means.

[1693] The "automatic countermeasure generation means" is a means for automatically generating an optimal countermeasure for a specified object.

[1694] The "command transmission means" is a means for transmitting an instruction to execute the generated countermeasure to the terminal.

[1695] The "countermeasure execution means" is a means for executing a specific countermeasure based on a received command.

[1696] The "re-monitoring means" is a means for re-monitoring the situation after the countermeasure has been implemented and collecting video and audio data again.

[1697] "Emotion recognition means" refers to a means for analyzing and recognizing a user's emotions.

[1698] The "optimization method" is a method for optimizing the content and intensity of countermeasures based on the recognized user emotions.

[1699] This invention is a system that uses cameras and microphones to collect video and audio in areas where pests and vermin appear, analyzes the video and audio to identify obstacles and suspicious individuals, and automatically generates countermeasures. This system is combined with an emotion recognition engine that recognizes the user's emotions and optimizes the content and strength of the countermeasures.

[1700] System Program

[1701] The system consists of the following components:

[1702] 1. Terminal

[1703] It includes a camera for collecting video data and a microphone for collecting audio data, and the cameras and microphones are installed inside and outside the home.

[1704] 2. Server

[1705] Data receiving means: Receives video and audio data transmitted from the terminal.

[1706] Data analysis method: Analyzes received video and audio data to identify objects (vermin, vermin, intruders, etc.). Face detection and abnormal sound detection models using OpenCV are used for the analysis.

[1707] Emotion recognition: Recognizes the user's emotions from the collected data and optimizes the strength and content of countermeasures based on that data. For emotion recognition, an emotion recognition model using Keras is used.

[1708] Countermeasure generation means: Automatically generates the optimal countermeasure based on the analysis results. For example, if an intruder is detected, a command to play an acoustic alarm is generated.

[1709] Command sending means: Sends the generated instructions for the solution to the terminal.

[1710] 3. Users

[1711] The system collects the user's emotional data (e.g., fear, surprise, etc.) and takes appropriate action based on that emotion.

[1712] Hardware and software used

[1713] Camera: Hardware for collecting video footage.

[1714] Microphone: Hardware for collecting sound.

[1715] Server: A central processing unit for data analysis and emotion recognition. A server with Python, OpenCV, and Keras installed is used.

[1716] Acoustic alarm device: Used as one of the countermeasures.

[1717] Specific examples of processing

[1718] For example, if an intruder appears near a home, the system will:

[1719] The devices (cameras and microphones installed around the house) capture video and audio of intruders and send the data to a server.

[1720] The server preprocesses the received video and audio data and uses an AI model to identify intruders. In parallel, an emotion recognition engine recognizes the user's emotions (e.g., fear or surprise).

[1721] The server generates an effective acoustic alarm against an intruder, and generates and sends to the terminal a command to adjust the intensity of the alarm according to the user's emotion data.

[1722] The device will play an audible alarm to scare off intruders, then continue monitoring to see if any further action is required.

[1723] Examples of prompt statements

[1724] "Design a smartphone application that runs inside the home. It uses a camera and microphone to capture video and audio in real time to detect intruders and suspicious activity. It uses an emotion recognition engine to detect user emotions (especially fear and surprise), and if an abnormality is detected, it will sound an alarm and notify the user or security company. Using Python, OpenCV, and Keras, you will pre-train the emotion recognition model and perform real-time analysis."

[1725] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1726] Step 1:

[1727] The device uses a camera and microphone to collect video and audio data in real time. This data includes the situation within the area, the user's facial expressions, and their comments. The collected data serves as basic information for understanding the situation within the area.

[1728] Step 2:

[1729] The terminals transmit the collected video and audio data to a server via a network. This transmission is performed in real time, and it is important to minimize delays.

[1730] Step 3:

[1731] The server preprocesses the received video and audio data, specifically removing noise and normalizing the data. This preprocessing improves data quality and increases analysis accuracy.

[1732] Step 4:

[1733] The server analyzes the preprocessed data. It uses an AI model to identify objects (such as suspicious people or pests) from the video data, and detects abnormal sounds (such as the sound of destruction or metal) from the audio data. The analysis results are used to confirm the presence of the object.

[1734] Step 5:

[1735] The server uses an emotion recognition engine to analyze the user's emotions. By analyzing facial expressions from video data and changes in tone and volume from audio data, it can determine what emotions the user is feeling (for example, fear or surprise).

[1736] Step 6:

[1737] The server generates the optimal countermeasure based on the analysis results. Specifically, it automatically determines the best action to take against the target object (for example, playing a loud cracking sound or emitting ultrasound). It also adjusts the strength and content of the countermeasure according to the user's emotions.

[1738] Step 7:

[1739] The server sends the generated instructions (commands) for the countermeasures to the terminal. These commands include the countermeasures to be executed and their parameters (e.g., type and intensity of sound, light pattern, etc.).

[1740] Step 8:

[1741] The device will then take action based on the command it receives. For example, if it receives a voice command, it will play a repellent sound from a specific speaker, and if it receives a video command, it will emit a strong light. This will allow it to take action against the target.

[1742] Step 9:

[1743] Even after the countermeasure is implemented, the device continues to monitor the situation within the area, collecting video and audio data again and sending it to the server.

[1744] Step 10:

[1745] The server analyzes the data again to see how the user's emotions have changed. Based on the analysis results, it determines whether further action is necessary. If the object still exists and further action is necessary, the process repeats from step 6.

[1746] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1747] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1748] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1749] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1750] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1751] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1752] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1753] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1754] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1755] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1756] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1757] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1758] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1759] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1760] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1761] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1762] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1763] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1764] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1765] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1766] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1767] The following is further disclosed regarding the above embodiment.

[1768] Drafting of claims

[1769] (Claim 1)

[1770] A means for collecting video and audio using a camera and a microphone installed in an area where pests or vermin appear;

[1771] means for receiving the collected video and audio data;

[1772] means for analyzing the received data and identifying the object;

[1773] A means for automatically generating a countermeasure according to the identified object;

[1774] means for sending a command to execute the generated remedy;

[1775] means for performing a remedial action based on the received command;

[1776] A means of re-monitoring and analyzing the situation after execution;

[1777] A system including:

[1778] (Claim 2)

[1779] 10. The system of claim 1, further comprising means for pre-processing the video and audio data to remove noise and normalize the data.

[1780] (Claim 3)

[1781] The system according to claim 1, further comprising means for performing data analysis using an AI model to identify an object.

[1782] (Claim 4)

[1783] 2. The system according to claim 1, further comprising means for automatically generating a response using appropriate sound, light, or image depending on the object.

[1784] (Claim 5)

[1785] 2. The system according to claim 1, further comprising means for collecting the status after the countermeasure is taken again and re-executing the countermeasure as necessary.

[1786] Above are the draft patent claims based on the system described.

[1787] "Example 1"

[1788] (Claim 1)

[1789] a means for collecting video and audio using an image capturing device and an audio recording device installed in an area where pests or vermin appear;

[1790] means for receiving the collected video and audio data;

[1791] means for analyzing the received data and identifying the object;

[1792] A means for automatically generating a countermeasure according to the identified object;

[1793] means for sending a command to execute the generated remedy;

[1794] means for performing a remedial action based on the received command;

[1795] A means of re-monitoring and analyzing the situation after execution;

[1796] A system including:

[1797] (Claim 2)

[1798] 10. The system of claim 1, further comprising means for pre-processing the video and audio data to remove noise and normalize the data.

[1799] (Claim 3)

[1800] The system of claim 1, further comprising means for performing data analysis using a generative AI model to identify an object.

[1801] "Application Example 1"

[1802] (Claim 1)

[1803] a means for collecting images and sounds using an imaging device and an audio collecting device installed in an area where pests or vermin appear;

[1804] means for transmitting the collected video and audio data to a mobile communication terminal;

[1805] A means for analyzing data received by the mobile communication terminal and identifying a target object;

[1806] means for automatically generating a countermeasure corresponding to the identified target object;

[1807] means for sending instructions to execute the generated countermeasure;

[1808] means for performing a remedial action based on the received command;

[1809] A means of re-monitoring and analyzing the situation after execution;

[1810] A system including:

[1811] (Claim 2)

[1812] 10. The system of claim 1, further comprising means for pre-processing the video and audio data to remove noise and normalize the data.

[1813] (Claim 3)

[1814] The system of claim 1, further comprising means for performing data analysis using a generative AI model to identify targets.

[1815] "Example 2: Combining Emotion Engines"

[1816] (Claim 1)

[1817] A means for collecting video and audio using a camera and a microphone installed in an area where pests or vermin appear;

[1818] means for receiving the collected video and audio data;

[1819] means for analyzing the received data and identifying the object;

[1820] means for recognizing a user's emotion;

[1821] A means for automatically generating a countermeasure according to the identified object and the user's emotion;

[1822] means for sending a command to execute the generated remedy;

[1823] means for performing a remedial action based on the received command;

[1824] means for re-monitoring and analyzing the situation and the user's emotions after execution;

[1825] A system including:

[1826] (Claim 2)

[1827] 10. The system of claim 1, further comprising means for pre-processing the video and audio data to remove noise and normalize the data.

[1828] (Claim 3)

[1829] The system of claim 1, further comprising means for performing data analysis using a generative AI model to identify objects and recognize user emotions.

[1830] "Application example 2 when combining emotion engines"

[1831] (Claim 1)

[1832] A means for collecting video and audio using a camera and a microphone installed in an area where pests or vermin appear;

[1833] means for receiving the collected video and audio data;

[1834] means for analyzing the received data and identifying the object;

[1835] A means for automatically generating a countermeasure according to the identified object;

[1836] means for sending a command to execute the generated remedy;

[1837] means for performing a remedial action based on the received command;

[1838] A means of re-monitoring and analyzing the situation after execution;

[1839] means for recognizing the user's emotions and optimizing the content and intensity of a countermeasure based on the emotions;

[1840] A system including:

[1841] (Claim 2)

[1842] 10. The system of claim 1, further comprising means for pre-processing the video and audio data to remove noise and normalize the data.

[1843] (Claim 3)

[1844] The system according to claim 1, further comprising means for performing data analysis using an AI model to identify an object. [Explanation of symbols]

[1845] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for collecting video and audio using a camera and a microphone installed in an area where pests or vermin appear; means for receiving the collected video and audio data; means for analyzing the received data and identifying the object; A means for automatically generating a countermeasure according to the identified object; means for sending a command to execute the generated remedy; means for performing a remedial action based on the received command; A means of re-monitoring and analyzing the situation after execution; A system including:

2. 2. The system of claim 1, further comprising means for pre-processing the video and audio data to remove noise and normalize the data.

3. The system according to claim 1, further comprising means for performing data analysis using an AI model to identify an object.

4. The system according to claim 1, further comprising means for automatically generating a countermeasure using appropriate sound, light, or image depending on the object.

5. 2. The system according to claim 1, further comprising means for collecting the status after the countermeasure is taken again and re-executing the countermeasure as necessary.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A