system

The system addresses privacy concerns in surveillance by converting personal features into abstract shapes using AI, enabling real-time privacy-protected surveillance.

JP2026071542APending Publication Date: 2026-04-30SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-17
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Conventional surveillance systems risk privacy violations due to direct recording and processing of video data, which can lead to data leakage and unauthorized access of personal information, necessitating a system that processes video in a form that does not allow individuals to be identified.

Method used

A system comprising a sensor for acquiring video, a computing means for real-time processing to abstract personal features into unidentifiable shapes, and a display means for presenting privacy-protected video, utilizing AI technology to convert individuals into abstract shapes.

Benefits of technology

Enables surveillance while protecting privacy by converting personal features into unidentifiable shapes in real time, ensuring privacy protection and efficient processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071542000001_ABST
    Figure 2026071542000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A sensor means for acquiring video, A calculation means for processing the video acquired by the sensor means, The calculation means generates an abstracted image, A display means for displaying the video generated by the generation means, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of this disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern monitoring systems, video data often contains personal privacy information, which causes concerns about privacy protection. In the conventional technology, since processing such as mosaicing is performed after all video data is recorded, there is a risk of leakage of privacy information due to data leakage or unauthorized access. Therefore, there is a need for a technology that processes video in a form that does not allow individuals to be identified from the time of shooting.

Means for Solving the Problems

[0005] The present invention provides a system comprising a sensor means for acquiring video, a computing means for processing the acquired video, a generation means for displaying the abstracted video generated by the computing means, and a display means, thereby enabling surveillance and observation while protecting privacy. The generation means converts the characteristics of a person into shapes that cannot be identified, and the computing means performs processing to abstract the video in real time. This makes it possible to generate video that does not contain any privacy information.

[0006] "Sensor means" refers to devices or modules used to acquire images, and typically refers to devices that use camera sensors.

[0007] "Computational means" refers to a processor or computer system that processes acquired video data, and is a device equipped with the function to perform complex calculations and image processing.

[0008] "Generation means" refers to hardware or software for generating new images based on data processed by computation means, and specifically refers to a system for outputting abstracted images.

[0009] "Display means" refers to a device for visually presenting images created by the generation means to the user, and includes display monitors. [Brief explanation of the drawing]

[0010] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5]This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0011] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0012] First, let's explain the terminology used in the following explanation.

[0013] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0014] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0015] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0016] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0017] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0018] [First Embodiment]

[0019] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0020] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0021] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0022] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0023] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0024] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0025] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0026] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0027] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0028] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0029] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0030] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0031] As an embodiment for carrying out the present invention, a surveillance system for the purpose of protecting privacy will be described. This system acquires images of the environment using a camera, which is a sensor means installed at a terminal, and processes the images in real time using a computing means on a server. The computing means uses an advanced image processing algorithm to abstract people in the images and convert them into shapes that make it impossible to identify individuals. As a result of this conversion process, images that protect privacy are output by a generation means.

[0032] As a concrete example, let's describe a surveillance camera system installed in a public facility. The camera, acting as the terminal, continuously acquires video footage from the location. This data is transmitted to a server via a network. The server's computing unit processes the incoming video data in real time, using AI technology to abstract the appearance of people. The generation unit generates the processed video based on this abstraction, and the display unit projects the abstracted video onto a monitor connected to the terminal.

[0033] Users can view privacy-protected video footage through their device's monitor. While the video allows for the identification of human movement and location, individual features are abstracted, making it impossible to identify specific individuals. Furthermore, this system is adaptable to diverse environments and is effective for surveillance in commercial facilities, hospitals, and educational institutions. Additionally, since the video is processed in real time, it can handle situations requiring immediate monitoring.

[0034] The following describes the processing flow.

[0035] Step 1:

[0036] The device activates its camera sensor to acquire video from the environment, capturing surrounding visual data in real time. The camera processes the video data frame by frame and stores it in an internal buffer.

[0037] Step 2:

[0038] The terminal transmits the acquired video data to the server via the network. Transmission takes place in real time, and the video data is encoded in an appropriate compression format (e.g., H.264) to ensure communication efficiency.

[0039] Step 3:

[0040] The server decodes the video data received from the terminal and prepares it for processing as individual frames. The decoded data is then converted into a format suitable for the image processing algorithm.

[0041] Step 4:

[0042] The server processes the video data frame by frame using an AI model. The AI ​​recognizes the shape and movement of people and converts them into unidentifiable silhouettes or abstract shapes.

[0043] Step 5:

[0044] The server re-encodes the abstracted video frames processed by the AI, generates them in real time using a generation mechanism, and sends them to the terminal. The network is optimized to ensure that the transmitted data arrives with low latency.

[0045] Step 6:

[0046] The terminal decodes abstracted video transmitted from the server and displays it in real time on a connected display device. This allows the user to monitor a privacy-protected environment.

[0047] Step 7:

[0048] Users can monitor the video displayed on their device to check for any unusual behavior or situations. They can also replay saved footage for further review if necessary. This allows users to use the camera system with peace of mind.

[0049] (Example 1)

[0050] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0051] In recent years, the need for privacy protection in surveillance systems has increased. However, conventional surveillance systems record individuals' appearances directly, which can lead to privacy violations. In addition, while surveillance data needs to be processed in real time, problems with processing delays and accuracy arise. To solve these problems, it is necessary to build a system that records data in a way that makes it impossible to identify individuals and processes it quickly.

[0052] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0053] In this invention, the server includes a measuring device that acquires video, an information processing device that processes the video acquired by the measuring device in real time and extracts characteristic points of a person, and a reconstruction device that generates an abstracted video from the information processing device. This makes it possible to efficiently process the video in real time while converting the appearance of an individual in the video into an unidentifiable shape.

[0054] A "measuring device" is a device used to acquire images of the environment and has the role of transmitting the image information to other devices.

[0055] An "information processing device" is a device that processes video data acquired from a measuring device in real time and extracts characteristic points of people in the video.

[0056] A "reconstruction device" is a device that generates abstract images that do not allow for the identification of individuals, based on feature points extracted by an information processing device.

[0057] A "monitoring system" refers to an entire system that combines measuring devices, information processing devices, and reconfiguration devices to monitor the environment while protecting privacy.

[0058] A "generative AI model" is a model that utilizes artificial intelligence technology used in information processing devices to achieve abstraction of video data.

[0059] This invention relates to a system for monitoring the environment while protecting privacy. The system mainly consists of a measuring device, an information processing device, and a reconstruction device, which are combined to implement the system.

[0060] The terminal unit's measuring device includes a camera for acquiring high-resolution images of the environment, and acquires video data in real time. This data is transmitted to a server via the network.

[0061] The server processes the received video data in real time using an information processing device. During this process, it utilizes a generative AI model to extract characteristic features of individuals and transforms these features into shapes that cannot identify individuals. The information processing device has advanced parallel processing capabilities and is designed to minimize processing delays.

[0062] The reconstruction device generates a new image based on abstracted feature points. This generated image is output to a monitor via a display device connected to the terminal, in a privacy-protected manner. The user can monitor through this privacy-protected image and take action as needed.

[0063] For example, in commercial facilities, the terminal's camera can capture the flow of customers, which can be used to understand congestion levels and consider safety measures. This system allows users, who are managers of commercial facilities, to understand the movements within the facility while ensuring the privacy of customers.

[0064] An example of a prompt message is: "Develop a program that abstracts individual characteristics for privacy protection and displays the video in real time."

[0065] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0066] Step 1:

[0067] The terminal's measuring device acquires video of the environment. It captures high-resolution video data in real time using an optical sensor as input. Specifically, it automatically adjusts exposure and focus according to ambient lighting conditions to improve data acquisition accuracy. The output is a stream of acquired video data, ready to be sent to the next processing step.

[0068] Step 2:

[0069] The terminal transmits the acquired video data to the server via the network. The input is the video data acquired in step 1. Specifically, the data is compressed and packetized, and stable data transfer is achieved through the communication protocol. As output, the data is transmitted to the server within the expected bandwidth.

[0070] Step 3:

[0071] The server's information processing unit receives the acquired video data in real time and begins processing it. The input is video data transmitted over the network. Using a generative AI model, it performs data calculations to extract feature points of people from the video, and based on this information, generates abstract data that makes it impossible to identify individuals. The output is video data abstracted based on the feature points.

[0072] Step 4:

[0073] The server's reconstruction device generates new video based on the abstracted video data. The input is the abstracted data obtained in step 3. Specifically, it applies mosaic and blurring processes to reconstruct the video in a state where people cannot be identified. The output is privacy-protected video data that can be used by display devices.

[0074] Step 5:

[0075] The terminal displays privacy-protected video footage sent from the server on its monitor. The input is video data generated by a reconstruction device. The specific operation involves adjusting the resolution and color correction to display the acquired video in a clear manner. The output appears on the monitor as an abstracted image that the user can view.

[0076] Step 6:

[0077] The user monitors the privacy-protected video displayed on the monitor. The input is the video displayed in step 5. Specific actions include rewinding the video as needed and focusing on specific areas to analyze movement patterns. The output is the decision based on the monitoring and the plan for implementing safety measures.

[0078] (Application Example 1)

[0079] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0080] With the widespread use of devices that acquire visual information, protecting privacy in public spaces has become increasingly important. However, conventional methods often display the visual characteristics of others in an identifiable way, making it difficult to completely protect individual privacy. Therefore, there is a need for a system that converts visual information in public spaces into a form that does not allow for personal identification, thereby enabling safe and secure information verification.

[0081] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0082] In this invention, the server includes detection means for extracting visual information, calculation means for transforming the visual information acquired by the detection means, and generation means for displaying the visual information generated by the calculation means. This makes it possible to abstract the visual information and confirm the situation in public spaces while preventing the identification of individuals.

[0083] "Visual information" refers to image and video data acquired by cameras and other light detection devices.

[0084] "Detection means" refers to sensors or devices used to collect visual information.

[0085] "Computational means" refers to algorithms and devices used to process acquired visual information and perform transformations and analyses as needed.

[0086] "Generating means" refers to devices or processes for creating a new form from visual information processed by computation means.

[0087] "Display means" refers to screens or display devices used to visually convey generated visual information to the user.

[0088] A "connection device" is an interface or network device used to exchange information with other devices or systems.

[0089] This invention provides a system that can effectively acquire visual information and confirm a situation while preventing the identification of personal information. To carry out the invention, a system including the following components is required.

[0090] The server uses cameras and other light detection devices as detection means to extract visual information. These detection means are placed in public spaces and acquire video in real time. The acquired video data is transmitted to the server via the network. Within the server, computational means abstract the visual information in a way that does not allow for individual identification. Specifically, software such as OpenCV and TENSORFLOW® are used to abstract facial features and remove personally identifiable information from the data.

[0091] The converted visual information is processed by a generation device. This processed information is presented by display devices such as smart glasses or display devices, ensuring privacy in public spaces. Through the connected device, users can view the abstracted visual information and effectively understand the situation in public places while protecting their privacy.

[0092] A concrete example is a security monitoring system in a commercial facility. Cameras installed within the facility acquire visual information, and a server processes this in real time to abstract its features, enabling monitoring of the facility while thoroughly protecting privacy. Furthermore, prompts such as "Design an AI algorithm that recognizes and anonymizes faces from video footage taken in the city" can be used to help effectively build generative AI models.

[0093] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0094] Step 1:

[0095] The device acquires visual information in real time using a camera. In this case, the input is the surrounding video, and the output is raw video data. The camera sensor captures the video and stores it as high-resolution image data.

[0096] Step 2:

[0097] The terminal transmits raw video data to the server over the network. The input is the raw video data acquired by the terminal, and the output is the video data received by the server. A communication protocol is used to ensure that the data is transmitted accurately.

[0098] Step 3:

[0099] The server starts a computing system to process the received video data. The input is the video data received by the server, and the output is abstracted video data. Here, face recognition is performed using OpenCV, and features are abstracted using TensorFlow. Features necessary for identifying individuals are abstracted industrially, preventing individual identification.

[0100] Step 4:

[0101] The server generates abstracted images based on a generative AI model. The input is abstracted video data, and the output is the generated visual information. The model transforms the data into avatars or mosaics, representing it in a way that protects privacy.

[0102] Step 5:

[0103] The server sends the generated visual information back to the terminal. The input is the generated visual information, and the output is the abstracted image displayed on the user's terminal. The transmission protocol ensures that the data is transferred in the correct format.

[0104] Step 6:

[0105] The user perceives abstracted visual information through the device's display mechanism. The input is the abstracted image displayed on the device, while the output is the visual information the user perceives visually. The display device uses a high-resolution screen to visually represent detailed movements and situations.

[0106] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0107] This invention provides a system combining sensor means, computation means, generation means, display means, and an emotion engine, designed to protect user privacy and recognize emotions. This system acquires images of the environment using a camera installed on a terminal. The acquired image data is transmitted in real time to a server's computation means and processed by an emotion engine, including AI.

[0108] The server's computational means abstracts the people in the video and converts them into unidentifiable shapes, while simultaneously recognizing the user's emotions using an emotion engine. The emotion engine analyzes subtle facial movements and identifies the emotional state from the user's expressions and actions. Based on this emotional information, the generation means can adjust the video to suit the user.

[0109] As a concrete example, consider a security system installed at a company's reception area. The reception camera, which acts as the terminal, captures the visitor's image, but for security reasons, the video is always abstracted and transformed into a form that makes it impossible to identify an individual. Furthermore, an emotion engine analyzes the visitor's facial expressions and infers emotions such as interest or anxiety. If the system detects anxiety in the visitor, it notifies the reception staff, enabling a quick response. In this way, this system can understand the visitor's emotions while protecting individual privacy and improving the user experience.

[0110] Users can not only view real-time, privacy-protected video through their device's monitor, but also obtain additional information through emotion recognition. This system is effective in enhancing security and service quality in various environments such as commercial facilities, hospitals, and educational institutions.

[0111] The following describes the processing flow.

[0112] Step 1:

[0113] The device activates its camera, which is a sensor, and captures the surrounding image in real time. The camera captures the image data frame by frame and stores it in an internal buffer.

[0114] Step 2:

[0115] The terminal compresses the acquired video data and sends it to the server over the network. The data is transmitted using an appropriate protocol (e.g., RTSP) to ensure low latency.

[0116] Step 3:

[0117] The server decodes the video data received from the terminal and prepares it for processing as individual frames. The decoded frames are then formatted appropriately for AI processing and sentiment analysis.

[0118] Step 4:

[0119] The server's processing system analyzes video frames using image processing algorithms, abstracting them into shapes that make it impossible to identify individuals. This process is for privacy protection, ensuring that individuals cannot be identified.

[0120] Step 5:

[0121] The server further analyzes the user's emotions from abstracted video frames using an emotion engine. It detects subtle facial movements and estimates emotions such as joy, anger, sadness, and happiness based on them.

[0122] Step 6:

[0123] The server generation method incorporates the results of emotion analysis into abstracted images and adjusts the image content as needed. This adjustment aims for optimal display in specific emotional states.

[0124] Step 7:

[0125] The generated video is compressed again and sent from the server to the terminal. The terminal decodes it and displays it on the monitor. The user can view this video and utilize additional information obtained through emotion recognition while monitoring a privacy-protected environment.

[0126] Step 8:

[0127] Users monitor the video displayed on the screen in real time, checking for any abnormalities and emotional states as needed. If specific emotional actions are required, users can respond quickly.

[0128] (Example 2)

[0129] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0130] Conventional video processing systems have struggled to simultaneously protect privacy and recognize user emotions. While abstracting video footage from surveillance cameras and sensors is necessary to protect individual privacy, this process inevitably leads to a decrease in the accuracy of emotion recognition. Furthermore, performing these processes in real time requires significant system processing power.

[0131] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0132] In this invention, the server includes detection means for acquiring video, computation means for processing and abstracting the video in real time, and analysis means for performing emotion recognition. This makes it possible to accurately recognize the user's emotions in real time while protecting the user's privacy.

[0133] "Detection means" refers to devices or sensors used to acquire video data from the environment.

[0134] "Computation means" refers to a device or process for processing acquired video data and performing abstraction as necessary.

[0135] "Production means" refers to devices or methods for generating processed and abstracted video data for subsequent processing or display.

[0136] "Presentation means" refers to a device or interface for visually providing generated video data to the user.

[0137] "Analysis means" refers to a device or method that recognizes a person's emotions based on acquired and processed video data and outputs the analysis results.

[0138] This invention provides a system for emotion recognition while protecting user privacy. The system primarily functions by acquiring video footage from a camera installed on a terminal and processing that data on a server. The following describes a specific implementation of this system.

[0139] Terminal role:

[0140] The device is equipped with a high-sensitivity camera for acquiring video. This camera can capture the environment and user's movements in real time. The device also has a communication module for transferring the acquired video data to a server, and the data is transmitted using a secure protocol.

[0141] Server role:

[0142] The server is equipped with high-performance computing devices to process the received video data. These devices utilize AI models as computational tools to perform data abstraction and emotion analysis. Video abstraction is essential from a privacy perspective, specifically transforming the images into shapes that make it impossible to identify individuals. Furthermore, the emotion engine has the ability to identify emotional states from the user's facial expressions and movements.

[0143] User roles:

[0144] Users can view privacy-protected video and analyzed sentiment information through their device's monitor. This enables users to respond appropriately and provide high-quality service to visitors.

[0145] Specific example:

[0146] For example, this system could be implemented in a company's reception area. In this case, the visitor's image would always be abstracted, making it impossible to identify an individual. On the other hand, the emotion engine would analyze the visitor's facial expressions and determine emotions such as interest or anxiety. If the system determines that a visitor is showing anxiety, the reception staff could respond to the visitor quickly.

[0147] Example of a prompt:

[0148] "Please describe in detail how you receive visitor facial expression data and analyze emotions using your emotion engine. Also, please describe how the system uses this analysis to improve the user experience while protecting privacy."

[0149] This system enables the provision of high-quality services while maintaining privacy in commercial facilities, medical institutions, educational institutions, and other settings.

[0150] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0151] Step 1:

[0152] The device uses a camera to acquire video footage of the environment in real time. At this point, the input is visual information of the surroundings. The device converts this video data into a digital format and sends it to the server in real time. Specifically, the camera continuously captures frames and transfers them to the server via a secure communication protocol.

[0153] Step 2:

[0154] The server inputs the received video data into a computing device. Here, initial data processing is performed to analyze the video data and extract important information. This process involves adjusting the video resolution and performing pre-processing such as noise reduction. The output is processed, clear video data.

[0155] Step 3:

[0156] The server's computing power abstracts the processed video data. Specifically, it uses machine learning algorithms to detect people in the video and transforms their features into shapes that are unidentifiable. The input for this step is preprocessed video data, and the output is abstracted video. The system achieves privacy protection at this stage.

[0157] Step 4:

[0158] The server inputs abstracted video into its emotion engine and performs detailed emotion analysis. The analysis method uses a machine learning model to infer emotions based on subtle facial changes in the video. The input for this step is abstracted video, and the output is identified emotion information. Specific operations include facial feature point extraction and emotional state classification.

[0159] Step 5:

[0160] The server inputs the extracted emotional information into a production mechanism that generates content for the user. This mechanism adjusts the images and information as needed based on the emotional information and presents it in a way that is useful to the user. The output of this step is information and images customized for the user. Specifically, this involves adjusting the screen display and creating notifications according to the emotional state.

[0161] Step 6:

[0162] Users view protected video and emotional information in real time through the presentation tools. They use this information to respond appropriately to visitors. The input in this step is customized information, and the output is the user's actions. Specific actions include the user providing situation-appropriate customer service and responses.

[0163] (Application Example 2)

[0164] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0165] In conventional stores, it has been difficult to analyze customers' facial expressions and behavior in real time and flexibly adjust customer service based on that analysis. Furthermore, it has been challenging to identify customer emotions while protecting privacy during video processing. This invention aims to solve these problems and realize efficient and privacy-conscious customer service that responds to customer emotions.

[0166] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0167] In this invention, the server includes detection means for acquiring video, computation means for processing the video acquired by the detection means, and creation means for generating an abstracted video by the computation means. This makes it possible to analyze customer emotions based on video that has been transformed into shapes that do not allow for the identification of individuals, and to present that information in real time.

[0168] "Detection means for acquiring images" refers to a sensor device for acquiring visual information of the physical environment.

[0169] "Computational means for processing the video acquired by the detection means" refers to a computer system for analyzing the acquired video data and extracting specific information.

[0170] "The calculation means that generates an abstracted image" refers to a device that provides a process for processing an image from data analyzed by the calculation means in a way that prevents the identification of individuals.

[0171] "Presentation means for displaying the image generated by the creation means" refers to a device having a display function that makes the abstracted image viewable by the user or system.

[0172] "An emotion analysis tool for analyzing a customer's emotional state" refers to an algorithm or system that identifies and determines an emotion from the facial expressions and actions of a person in a video.

[0173] "Information presentation means for displaying emotional information analyzed by the emotional analysis means" refers to a device or system for providing information obtained through emotional analysis to a user visually or by other means.

[0174] The system of this invention is built to improve the customer experience in physical stores. The system includes the following hardware and software configuration.

[0175] Hardware configuration:

[0176] Detection devices capable of acquiring information at high resolution (e.g., cameras)

[0177] Visual presentation devices equipped with video output displays (e.g., smart glasses)

[0178] Software configuration:

[0179] A computational program for image processing (e.g., OpenCV library)

[0180] An emotion analysis processor that analyzes emotions from a customer's facial expressions (e.g., Microsoft® Azure® Emotion API)

[0181] The server first acquires images of customers using detection devices installed in the store. Then, it processes the acquired video data using computational means, abstracting it in a way that protects individual privacy while preventing identification. Next, an emotion analysis means analyzes the customer's emotional state from the abstracted video, and the results are displayed in real time on a visual display device worn by staff using an information display means. This allows staff to instantly understand the customer's interests and anxieties, enabling them to take more appropriate action.

[0182] For example, when a new product is introduced in a store, if a customer is looking at the new product with interest, the visual display device will show an analysis result indicating "showing interest." This allows staff to proactively initiate customer service, explaining and suggesting the product.

[0183] Examples of prompts include, "Analyze the facial expressions of customers in the store and analyze their emotions, such as interest and anxiety, in real time," and "Based on the changes in customer emotions, suggest appropriate countermeasures to the service staff."

[0184] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0185] Step 1:

[0186] The server acquires video from detection devices installed in physical stores. The input is real-time camera footage, and the output is saved as video data. The server immediately prepares to place this video data into a processing queue.

[0187] Step 2:

[0188] The server analyzes the acquired video using computational means. The input is the video data acquired in step 1, and the output is abstracted video data. The server uses an image processing library (e.g., OpenCV) to abstract people in the video and convert them into shapes that do not allow for individual identification.

[0189] Step 3:

[0190] The server feeds abstracted video data into an emotion analysis system to analyze the customer's emotional state. The input is the abstracted data generated in step 2, and the output is the analyzed emotional information. A generative AI model is used to identify emotions from the abstracted facial features and extract those states as numerical data or labels.

[0191] Step 4:

[0192] The server transmits the analyzed emotion information to a visual display device, and the user provides customer service based on the information displayed on the device. The input is the emotion information obtained in step 3, and the output is the emotion label or numerical information displayed on the visual display device. Based on this, the user considers appropriate suggestions and customer service strategies for the customer.

[0193] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0194] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0195] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0196] [Second Embodiment]

[0197] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0198] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0199] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0200] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0201] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0202] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0203] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0204] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0205] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0206] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0207] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0208] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0209] As an embodiment for carrying out the present invention, a surveillance system for the purpose of protecting privacy will be described. This system acquires images of the environment using a camera, which is a sensor means installed at a terminal, and processes the images in real time using a computing means on a server. The computing means uses an advanced image processing algorithm to abstract people in the images and convert them into shapes that make it impossible to identify individuals. As a result of this conversion process, images that protect privacy are output by a generation means.

[0210] As a concrete example, let's describe a surveillance camera system installed in a public facility. The camera, acting as the terminal, continuously acquires video footage from the location. This data is transmitted to a server via a network. The server's computing unit processes the incoming video data in real time, using AI technology to abstract the appearance of people. The generation unit generates the processed video based on this abstraction, and the display unit projects the abstracted video onto a monitor connected to the terminal.

[0211] Users can view privacy-protected video footage through their device's monitor. While the video allows for the identification of human movement and location, individual features are abstracted, making it impossible to identify specific individuals. Furthermore, this system is adaptable to diverse environments and is effective for surveillance in commercial facilities, hospitals, and educational institutions. Additionally, since the video is processed in real time, it can handle situations requiring immediate monitoring.

[0212] The following describes the processing flow.

[0213] Step 1:

[0214] The device activates its camera sensor to acquire video from the environment, capturing surrounding visual data in real time. The camera processes the video data frame by frame and stores it in an internal buffer.

[0215] Step 2:

[0216] The terminal transmits the acquired video data to the server via the network. Transmission takes place in real time, and the video data is encoded in an appropriate compression format (e.g., H.264) to ensure communication efficiency.

[0217] Step 3:

[0218] The server decodes the video data received from the terminal and prepares it for processing as individual frames. The decoded data is then converted into a format suitable for the image processing algorithm.

[0219] Step 4:

[0220] The server processes the video data frame by frame using an AI model. The AI ​​recognizes the shape and movement of people and converts them into unidentifiable silhouettes or abstract shapes.

[0221] Step 5:

[0222] The server re-encodes the abstracted video frames processed by the AI, generates them in real time using a generation mechanism, and sends them to the terminal. The network is optimized to ensure that the transmitted data arrives with low latency.

[0223] Step 6:

[0224] The terminal decodes abstracted video transmitted from the server and displays it in real time on a connected display device. This allows the user to monitor a privacy-protected environment.

[0225] Step 7:

[0226] Users can monitor the video displayed on their device to check for any unusual behavior or situations. They can also replay saved footage for further review if necessary. This allows users to use the camera system with peace of mind.

[0227] (Example 1)

[0228] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0229] In recent years, the need for privacy protection in surveillance systems has increased. However, conventional surveillance systems record individuals' appearances directly, which can lead to privacy violations. In addition, while surveillance data needs to be processed in real time, problems with processing delays and accuracy arise. To solve these problems, it is necessary to build a system that records data in a way that makes it impossible to identify individuals and processes it quickly.

[0230] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0231] In this invention, the server includes a measuring device that acquires video, an information processing device that processes the video acquired by the measuring device in real time and extracts characteristic points of a person, and a reconstruction device that generates an abstracted video from the information processing device. This makes it possible to efficiently process the video in real time while converting the appearance of an individual in the video into an unidentifiable shape.

[0232] A "measuring device" is a device used to acquire images of the environment and has the role of transmitting the image information to other devices.

[0233] An "information processing device" is a device that processes video data acquired from a measuring device in real time and extracts characteristic points of people in the video.

[0234] A "reconstruction device" is a device that generates abstract images that do not allow for the identification of individuals, based on feature points extracted by an information processing device.

[0235] A "monitoring system" refers to an entire system that combines measuring devices, information processing devices, and reconfiguration devices to monitor the environment while protecting privacy.

[0236] A "generative AI model" is a model that utilizes artificial intelligence technology used in information processing devices to achieve abstraction of video data.

[0237] This invention relates to a system for monitoring the environment while protecting privacy. The system mainly consists of a measuring device, an information processing device, and a reconstruction device, which are combined to implement the system.

[0238] The terminal unit's measuring device includes a camera for acquiring high-resolution images of the environment, and acquires video data in real time. This data is transmitted to a server via the network.

[0239] The server processes the received video data in real time using an information processing device. During this process, it utilizes a generative AI model to extract characteristic features of individuals and transforms these features into shapes that cannot identify individuals. The information processing device has advanced parallel processing capabilities and is designed to minimize processing delays.

[0240] The reconstruction device generates a new image based on abstracted feature points. This generated image is output to a monitor via a display device connected to the terminal, in a privacy-protected manner. The user can monitor through this privacy-protected image and take action as needed.

[0241] For example, in commercial facilities, the terminal's camera can capture the flow of customers, which can be used to understand congestion levels and consider safety measures. This system allows users, who are managers of commercial facilities, to understand the movements within the facility while ensuring the privacy of customers.

[0242] An example of a prompt message is: "Develop a program that abstracts individual characteristics for privacy protection and displays the video in real time."

[0243] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0244] Step 1:

[0245] The terminal's measuring device acquires video of the environment. It captures high-resolution video data in real time using an optical sensor as input. Specifically, it automatically adjusts exposure and focus according to ambient lighting conditions to improve data acquisition accuracy. The output is a stream of acquired video data, ready to be sent to the next processing step.

[0246] Step 2:

[0247] The terminal transmits the acquired video data to the server via the network. The input is the video data acquired in step 1. Specifically, the data is compressed and packetized, and stable data transfer is achieved through the communication protocol. As output, the data is transmitted to the server within the expected bandwidth.

[0248] Step 3:

[0249] The server's information processing unit receives the acquired video data in real time and begins processing it. The input is video data transmitted over the network. Using a generative AI model, it performs data calculations to extract feature points of people from the video, and based on this information, generates abstract data that makes it impossible to identify individuals. The output is video data abstracted based on the feature points.

[0250] Step 4:

[0251] The server's reconstruction device generates new video based on the abstracted video data. The input is the abstracted data obtained in step 3. Specifically, it applies mosaic and blurring processes to reconstruct the video in a state where people cannot be identified. The output is privacy-protected video data that can be used by display devices.

[0252] Step 5:

[0253] The terminal displays privacy-protected video footage sent from the server on its monitor. The input is video data generated by a reconstruction device. The specific operation involves adjusting the resolution and color correction to display the acquired video in a clear manner. The output appears on the monitor as an abstracted image that the user can view.

[0254] Step 6:

[0255] The user monitors the privacy-protected video displayed on the monitor. The input is the video displayed in step 5. Specific actions include rewinding the video as needed and focusing on specific areas to analyze movement patterns. The output is the decision based on the monitoring and the plan for implementing safety measures.

[0256] (Application Example 1)

[0257] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0258] With the widespread use of devices that acquire visual information, protecting privacy in public spaces has become increasingly important. However, conventional methods often display the visual characteristics of others in an identifiable way, making it difficult to completely protect individual privacy. Therefore, there is a need for a system that converts visual information in public spaces into a form that does not allow for personal identification, thereby enabling safe and secure information verification.

[0259] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0260] In this invention, the server includes detection means for extracting visual information, calculation means for transforming the visual information acquired by the detection means, and generation means for displaying the visual information generated by the calculation means. This makes it possible to abstract the visual information and confirm the situation in public spaces while preventing the identification of individuals.

[0261] "Visual information" refers to image and video data acquired by cameras and other light detection devices.

[0262] "Detection means" refers to sensors or devices used to collect visual information.

[0263] "Computational means" refers to algorithms and devices used to process acquired visual information and perform transformations and analyses as needed.

[0264] "Generating means" refers to devices or processes for creating a new form from visual information processed by computation means.

[0265] "Display means" refers to screens or display devices used to visually convey generated visual information to the user.

[0266] A "connection device" is an interface or network device used to exchange information with other devices or systems.

[0267] This invention provides a system that can effectively acquire visual information and confirm a situation while preventing the identification of personal information. To carry out the invention, a system including the following components is required.

[0268] The server uses cameras and other light detection devices as detection means to extract visual information. These detection means are placed in public spaces and acquire video in real time. The acquired video data is transmitted to the server via the network. Within the server, computational means abstract the visual information in a way that does not allow for individual identification. Specifically, software such as OpenCV and TensorFlow is used to abstract facial features and remove personally identifiable information from the data.

[0269] The converted visual information is processed by a generation device. This processed information is presented by display devices such as smart glasses or display devices, ensuring privacy in public spaces. Through the connected device, users can view the abstracted visual information and effectively understand the situation in public places while protecting their privacy.

[0270] A concrete example is a security monitoring system in a commercial facility. Cameras installed within the facility acquire visual information, and a server processes this in real time to abstract its features, enabling monitoring of the facility while thoroughly protecting privacy. Furthermore, prompts such as "Design an AI algorithm that recognizes and anonymizes faces from video footage taken in the city" can be used to help effectively build generative AI models.

[0271] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0272] Step 1:

[0273] The device acquires visual information in real time using a camera. In this case, the input is the surrounding video, and the output is raw video data. The camera sensor captures the video and stores it as high-resolution image data.

[0274] Step 2:

[0275] The terminal transmits raw video data to the server over the network. The input is the raw video data acquired by the terminal, and the output is the video data received by the server. A communication protocol is used to ensure that the data is transmitted accurately.

[0276] Step 3:

[0277] The server activates the computing means to process the received video data. The input is the video data received by the server, and the output is the abstracted video data. Here, face recognition is performed using OpenCV, and features are abstracted using TensorFlow. Industrially abstract the features necessary to identify an individual to prevent the identification of the individual.

[0278] Step 4:

[0279] The server forms the abstracted video by the generating means based on the generative AI model. The input is the abstracted video data, and the output is the generated visual information. The data is converted into an avatar or mosaic by the model and presented in a form that protects privacy.

[0280] Step 5:

[0281] The server returns the generated visual information to the terminal. The input is the generated visual information, and the output is the abstracted video displayed on the user's terminal. The data is transferred in the correct format according to the transmission protocol.

[0282] Step 6:

[0283] The user checks the abstracted visual information through the display means of the terminal. The input is the abstracted video displayed on the terminal, and the output is the video information visually recognized by the user. The display device uses a high-resolution screen to represent visually detailed movements and situations.

[0284] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion specific model 59 and perform specific processing using the user's emotion.

[0285] The present invention provides a system that combines sensor means, computing means, generation means, display means, and an emotion engine, which is designed to protect user privacy and recognize emotions. This system acquires video of the environment through a camera installed on a terminal. The acquired video data is transmitted in real-time to the computing means of the server and processed by an emotion engine that includes AI.

[0286] The computing means of the server abstracts the people in the video and converts them into an unidentifiable shape, while at the same time recognizing the user's emotions using the emotion engine. The emotion engine analyzes the subtle movements of the face and identifies the emotional state from the user's expressions and actions. Also, based on this emotion information, the generation means can adjust the video in a form suitable for the user.

[0287] As a specific example, consider a security system installed at the reception of a company. The reception camera, which is the terminal, captures the appearance of visitors, but from a security perspective, the video is always abstracted and transformed into a form where individuals cannot be identified. Furthermore, the emotion engine analyzes the expressions of the visitors and推测 emotions such as interest and uneasiness. When the system detects uneasiness in a visitor, the reception staff is notified, enabling a prompt response. In this way, this system can protect an individual's privacy while understanding the emotions of visitors and improving the user experience.

[0288] The user can not only view the video in a state where privacy is protected in real-time through the monitor of the terminal, but also obtain additional information through emotion recognition. This system is effective for enhancing security and service quality in various environments such as commercial facilities, hospitals, educational institutions, etc.

[0289] The following describes the processing flow.

[0290] Step 1:

[0291] The device activates its camera, which is a sensor, and captures the surrounding image in real time. The camera captures the image data frame by frame and stores it in an internal buffer.

[0292] Step 2:

[0293] The terminal compresses the acquired video data and sends it to the server over the network. The data is transmitted using an appropriate protocol (e.g., RTSP) to ensure low latency.

[0294] Step 3:

[0295] The server decodes the video data received from the terminal and prepares it for processing as individual frames. The decoded frames are then formatted appropriately for AI processing and sentiment analysis.

[0296] Step 4:

[0297] The server's processing system analyzes video frames using image processing algorithms, abstracting them into shapes that make it impossible to identify individuals. This process is for privacy protection, ensuring that individuals cannot be identified.

[0298] Step 5:

[0299] The server further analyzes the user's emotions from abstracted video frames using an emotion engine. It detects subtle facial movements and estimates emotions such as joy, anger, sadness, and happiness based on them.

[0300] Step 6:

[0301] The server generation method incorporates the results of emotion analysis into abstracted images and adjusts the image content as needed. This adjustment aims for optimal display in specific emotional states.

[0302] Step 7:

[0303] The generated video is compressed again and transmitted from the server to the terminal. The terminal decodes this and displays it on the monitor. The user can watch this video and utilize the additional information obtained by emotion recognition while monitoring the privacy-protected environment.

[0304] Step 8:

[0305] The user monitors the video displayed on the monitor in real time and checks for abnormalities or emotional states as needed. If actions based on specific emotions are required, the user can respond promptly.

[0306] (Example 2)

[0307] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0308] In a conventional video processing system, it has been difficult to perform privacy protection and user emotion recognition simultaneously. In video acquisition by surveillance cameras and sensors, it is necessary to abstract the video in order to protect an individual's privacy, but there has been a problem that the accuracy of emotion recognition decreases accordingly. Furthermore, in order to perform these processes in real time, the processing ability of the system is required.

[0309] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following respective means.

[0310] In this invention, the server includes a detection means for acquiring video, a calculation means for processing and abstracting the video in real time, and an analysis means for performing emotion recognition. Thereby, while protecting the privacy of the user, it becomes possible to accurately recognize the user's emotion in real time.

[0311] The "detection means" is a device or sensor for acquiring video data from the environment.

[0312] "Computation means" refers to a device or process for processing acquired video data and performing abstraction as necessary.

[0313] "Production means" refers to devices or methods for generating processed and abstracted video data for subsequent processing or display.

[0314] "Presentation means" refers to a device or interface for visually providing generated video data to the user.

[0315] "Analysis means" refers to a device or method that recognizes a person's emotions based on acquired and processed video data and outputs the analysis results.

[0316] This invention provides a system for emotion recognition while protecting user privacy. The system primarily functions by acquiring video footage from a camera installed on a terminal and processing that data on a server. The following describes a specific implementation of this system.

[0317] Terminal role:

[0318] The device is equipped with a high-sensitivity camera for acquiring video. This camera can capture the environment and user's movements in real time. The device also has a communication module for transferring the acquired video data to a server, and the data is transmitted using a secure protocol.

[0319] Server role:

[0320] The server is equipped with high-performance computing devices to process the received video data. These devices utilize AI models as computational tools to perform data abstraction and emotion analysis. Video abstraction is essential from a privacy perspective, specifically transforming the images into shapes that make it impossible to identify individuals. Furthermore, the emotion engine has the ability to identify emotional states from the user's facial expressions and movements.

[0321] User roles:

[0322] Users can view privacy-protected video and analyzed sentiment information through their device's monitor. This enables users to respond appropriately and provide high-quality service to visitors.

[0323] Specific example:

[0324] For example, this system could be implemented in a company's reception area. In this case, the visitor's image would always be abstracted, making it impossible to identify an individual. On the other hand, the emotion engine would analyze the visitor's facial expressions and determine emotions such as interest or anxiety. If the system determines that a visitor is showing anxiety, the reception staff could respond to the visitor quickly.

[0325] Example of a prompt:

[0326] "Please describe in detail how you receive visitor facial expression data and analyze emotions using your emotion engine. Also, please describe how the system uses this analysis to improve the user experience while protecting privacy."

[0327] This system enables the provision of high-quality services while maintaining privacy in commercial facilities, medical institutions, educational institutions, and other settings.

[0328] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0329] Step 1:

[0330] The device uses a camera to acquire video footage of the environment in real time. At this point, the input is visual information of the surroundings. The device converts this video data into a digital format and sends it to the server in real time. Specifically, the camera continuously captures frames and transfers them to the server via a secure communication protocol.

[0331] Step 2:

[0332] The server inputs the received video data into a computing device. Here, initial data processing is performed to analyze the video data and extract important information. This process involves adjusting the video resolution and performing pre-processing such as noise reduction. The output is processed, clear video data.

[0333] Step 3:

[0334] The server's computing power abstracts the processed video data. Specifically, it uses machine learning algorithms to detect people in the video and transforms their features into shapes that are unidentifiable. The input for this step is preprocessed video data, and the output is abstracted video. The system achieves privacy protection at this stage.

[0335] Step 4:

[0336] The server inputs abstracted video into its emotion engine and performs detailed emotion analysis. The analysis method uses a machine learning model to infer emotions based on subtle facial changes in the video. The input for this step is abstracted video, and the output is identified emotion information. Specific operations include facial feature point extraction and emotional state classification.

[0337] Step 5:

[0338] The server inputs the extracted emotional information into a production mechanism that generates content for the user. This mechanism adjusts the images and information as needed based on the emotional information and presents it in a way that is useful to the user. The output of this step is information and images customized for the user. Specifically, this involves adjusting the screen display and creating notifications according to the emotional state.

[0339] Step 6:

[0340] Users view protected video and emotional information in real time through the presentation tools. They use this information to respond appropriately to visitors. The input in this step is customized information, and the output is the user's actions. Specific actions include the user providing situation-appropriate customer service and responses.

[0341] (Application Example 2)

[0342] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0343] In conventional stores, it has been difficult to analyze customers' facial expressions and behavior in real time and flexibly adjust customer service based on that analysis. Furthermore, it has been challenging to identify customer emotions while protecting privacy during video processing. This invention aims to solve these problems and realize efficient and privacy-conscious customer service that responds to customer emotions.

[0344] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0345] In this invention, the server includes detection means for acquiring video, computation means for processing the video acquired by the detection means, and creation means for generating an abstracted video by the computation means. This makes it possible to analyze customer emotions based on video that has been transformed into shapes that do not allow for the identification of individuals, and to present that information in real time.

[0346] "Detection means for acquiring images" refers to a sensor device for acquiring visual information of the physical environment.

[0347] "Computational means for processing the video acquired by the detection means" refers to a computer system for analyzing the acquired video data and extracting specific information.

[0348] "The calculation means that generates an abstracted image" refers to a device that provides a process for processing an image from data analyzed by the calculation means in a way that prevents the identification of individuals.

[0349] "Presentation means for displaying the image generated by the creation means" refers to a device having a display function that makes the abstracted image viewable by the user or system.

[0350] "An emotion analysis tool for analyzing a customer's emotional state" refers to an algorithm or system that identifies and determines an emotion from the facial expressions and actions of a person in a video.

[0351] "Information presentation means for displaying emotional information analyzed by the emotional analysis means" refers to a device or system for providing information obtained through emotional analysis to a user visually or by other means.

[0352] The system of this invention is built to improve the customer experience in physical stores. The system includes the following hardware and software configuration.

[0353] Hardware configuration:

[0354] Detection devices capable of acquiring information at high resolution (e.g., cameras)

[0355] Visual presentation devices equipped with video output displays (e.g., smart glasses)

[0356] Software configuration:

[0357] A computational program for image processing (e.g., OpenCV library)

[0358] An emotion analysis processor that analyzes emotions from a customer's facial expressions (e.g., Microsoft Azure Emotion API)

[0359] The server first acquires images of customers using detection devices installed in the store. Then, it processes the acquired video data using computational means, abstracting it in a way that protects individual privacy while preventing identification. Next, an emotion analysis means analyzes the customer's emotional state from the abstracted video, and the results are displayed in real time on a visual display device worn by staff using an information display means. This allows staff to instantly understand the customer's interests and anxieties, enabling them to take more appropriate action.

[0360] For example, when a new product is introduced in a store, if a customer is looking at the new product with interest, the visual display device will show an analysis result indicating "showing interest." This allows staff to proactively initiate customer service, explaining and suggesting the product.

[0361] Examples of prompts include, "Analyze the facial expressions of customers in the store and analyze their emotions, such as interest and anxiety, in real time," and "Based on the changes in customer emotions, suggest appropriate countermeasures to the service staff."

[0362] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0363] Step 1:

[0364] The server acquires video from detection devices installed in physical stores. The input is real-time camera footage, and the output is saved as video data. The server immediately prepares to place this video data into a processing queue.

[0365] Step 2:

[0366] The server analyzes the acquired video using computational means. The input is the video data acquired in step 1, and the output is abstracted video data. The server uses an image processing library (e.g., OpenCV) to abstract people in the video and convert them into shapes that do not allow for individual identification.

[0367] Step 3:

[0368] The server feeds abstracted video data into an emotion analysis system to analyze the customer's emotional state. The input is the abstracted data generated in step 2, and the output is the analyzed emotional information. A generative AI model is used to identify emotions from the abstracted facial features and extract those states as numerical data or labels.

[0369] Step 4:

[0370] The server transmits the analyzed emotion information to a visual display device, and the user provides customer service based on the information displayed on the device. The input is the emotion information obtained in step 3, and the output is the emotion label or numerical information displayed on the visual display device. Based on this, the user considers appropriate suggestions and customer service strategies for the customer.

[0371] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0372] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0373] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0374] [Third Embodiment]

[0375] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0376] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0377] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0378] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0379] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0380] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0381] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0382] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0383] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0384] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0385] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0386] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0387] As an embodiment for carrying out the present invention, a surveillance system for the purpose of protecting privacy will be described. This system acquires images of the environment using a camera, which is a sensor means installed at a terminal, and processes the images in real time using a computing means on a server. The computing means uses an advanced image processing algorithm to abstract people in the images and convert them into shapes that make it impossible to identify individuals. As a result of this conversion process, images that protect privacy are output by a generation means.

[0388] As a concrete example, let's describe a surveillance camera system installed in a public facility. The camera, acting as the terminal, continuously acquires video footage from the location. This data is transmitted to a server via a network. The server's computing unit processes the incoming video data in real time, using AI technology to abstract the appearance of people. The generation unit generates the processed video based on this abstraction, and the display unit projects the abstracted video onto a monitor connected to the terminal.

[0389] Users can view privacy-protected video footage through their device's monitor. While the video allows for the identification of human movement and location, individual features are abstracted, making it impossible to identify specific individuals. Furthermore, this system is adaptable to diverse environments and is effective for surveillance in commercial facilities, hospitals, and educational institutions. Additionally, since the video is processed in real time, it can handle situations requiring immediate monitoring.

[0390] The following describes the processing flow.

[0391] Step 1:

[0392] The device activates its camera sensor to acquire video from the environment, capturing surrounding visual data in real time. The camera processes the video data frame by frame and stores it in an internal buffer.

[0393] Step 2:

[0394] The terminal transmits the acquired video data to the server via the network. Transmission takes place in real time, and the video data is encoded in an appropriate compression format (e.g., H.264) to ensure communication efficiency.

[0395] Step 3:

[0396] The server decodes the video data received from the terminal and prepares it for processing as individual frames. The decoded data is then converted into a format suitable for the image processing algorithm.

[0397] Step 4:

[0398] The server processes the video data frame by frame using an AI model. The AI ​​recognizes the shape and movement of people and converts them into unidentifiable silhouettes or abstract shapes.

[0399] Step 5:

[0400] The server re-encodes the abstracted video frames processed by the AI, generates them in real time using a generation mechanism, and sends them to the terminal. The network is optimized to ensure that the transmitted data arrives with low latency.

[0401] Step 6:

[0402] The terminal decodes abstracted video transmitted from the server and displays it in real time on a connected display device. This allows the user to monitor a privacy-protected environment.

[0403] Step 7:

[0404] Users can monitor the video displayed on their device to check for any unusual behavior or situations. They can also replay saved footage for further review if necessary. This allows users to use the camera system with peace of mind.

[0405] (Example 1)

[0406] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0407] In recent years, the need for privacy protection in surveillance systems has increased. However, conventional surveillance systems record individuals' appearances directly, which can lead to privacy violations. In addition, while surveillance data needs to be processed in real time, problems with processing delays and accuracy arise. To solve these problems, it is necessary to build a system that records data in a way that makes it impossible to identify individuals and processes it quickly.

[0408] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0409] In this invention, the server includes a measuring device that acquires video, an information processing device that processes the video acquired by the measuring device in real time and extracts characteristic points of a person, and a reconstruction device that generates an abstracted video from the information processing device. This makes it possible to efficiently process the video in real time while converting the appearance of an individual in the video into an unidentifiable shape.

[0410] A "measuring device" is a device used to acquire images of the environment and has the role of transmitting the image information to other devices.

[0411] An "information processing device" is a device that processes video data acquired from a measuring device in real time and extracts characteristic points of people in the video.

[0412] A "reconstruction device" is a device that generates abstract images that do not allow for the identification of individuals, based on feature points extracted by an information processing device.

[0413] A "monitoring system" refers to an entire system that combines measuring devices, information processing devices, and reconfiguration devices to monitor the environment while protecting privacy.

[0414] A "generative AI model" is a model that utilizes artificial intelligence technology used in information processing devices to achieve abstraction of video data.

[0415] This invention relates to a system for monitoring the environment while protecting privacy. The system mainly consists of a measuring device, an information processing device, and a reconstruction device, which are combined to implement the system.

[0416] The terminal unit's measuring device includes a camera for acquiring high-resolution images of the environment, and acquires video data in real time. This data is transmitted to a server via the network.

[0417] The server processes the received video data in real time using an information processing device. During this process, it utilizes a generative AI model to extract characteristic features of individuals and transforms these features into shapes that cannot identify individuals. The information processing device has advanced parallel processing capabilities and is designed to minimize processing delays.

[0418] The reconstruction device generates a new image based on abstracted feature points. This generated image is output to a monitor via a display device connected to the terminal, in a privacy-protected manner. The user can monitor through this privacy-protected image and take action as needed.

[0419] For example, in commercial facilities, the terminal's camera can capture the flow of customers, which can be used to understand congestion levels and consider safety measures. This system allows users, who are managers of commercial facilities, to understand the movements within the facility while ensuring the privacy of customers.

[0420] An example of a prompt message is: "Develop a program that abstracts individual characteristics for privacy protection and displays the video in real time."

[0421] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0422] Step 1:

[0423] The terminal's measuring device acquires video of the environment. It captures high-resolution video data in real time using an optical sensor as input. Specifically, it automatically adjusts exposure and focus according to ambient lighting conditions to improve data acquisition accuracy. The output is a stream of acquired video data, ready to be sent to the next processing step.

[0424] Step 2:

[0425] The terminal transmits the acquired video data to the server via the network. The input is the video data acquired in step 1. Specifically, the data is compressed and packetized, and stable data transfer is achieved through the communication protocol. As output, the data is transmitted to the server within the expected bandwidth.

[0426] Step 3:

[0427] The server's information processing unit receives the acquired video data in real time and begins processing it. The input is video data transmitted over the network. Using a generative AI model, it performs data calculations to extract feature points of people from the video, and based on this information, generates abstract data that makes it impossible to identify individuals. The output is video data abstracted based on the feature points.

[0428] Step 4:

[0429] The server's reconstruction device generates new video based on the abstracted video data. The input is the abstracted data obtained in step 3. Specifically, it applies mosaic and blurring processes to reconstruct the video in a state where people cannot be identified. The output is privacy-protected video data that can be used by display devices.

[0430] Step 5:

[0431] The terminal displays privacy-protected video footage sent from the server on its monitor. The input is video data generated by a reconstruction device. The specific operation involves adjusting the resolution and color correction to display the acquired video in a clear manner. The output appears on the monitor as an abstracted image that the user can view.

[0432] Step 6:

[0433] The user monitors the privacy-protected video displayed on the monitor. The input is the video displayed in step 5. Specific actions include rewinding the video as needed and focusing on specific areas to analyze movement patterns. The output is the decision based on the monitoring and the plan for implementing safety measures.

[0434] (Application Example 1)

[0435] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0436] With the widespread use of devices that acquire visual information, protecting privacy in public spaces has become increasingly important. However, conventional methods often display the visual characteristics of others in an identifiable way, making it difficult to completely protect individual privacy. Therefore, there is a need for a system that converts visual information in public spaces into a form that does not allow for personal identification, thereby enabling safe and secure information verification.

[0437] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0438] In this invention, the server includes detection means for extracting visual information, calculation means for transforming the visual information acquired by the detection means, and generation means for displaying the visual information generated by the calculation means. This makes it possible to abstract the visual information and confirm the situation in public spaces while preventing the identification of individuals.

[0439] "Visual information" refers to image and video data acquired by cameras and other light detection devices.

[0440] "Detection means" refers to sensors or devices used to collect visual information.

[0441] "Computational means" refers to algorithms and devices used to process acquired visual information and perform transformations and analyses as needed.

[0442] "Generating means" refers to devices or processes for creating a new form from visual information processed by computation means.

[0443] "Display means" refers to screens or display devices used to visually convey generated visual information to the user.

[0444] A "connection device" is an interface or network device used to exchange information with other devices or systems.

[0445] This invention provides a system that can effectively acquire visual information and confirm a situation while preventing the identification of personal information. To carry out the invention, a system including the following components is required.

[0446] The server uses cameras and other light detection devices as detection means to extract visual information. These detection means are placed in public spaces and acquire video in real time. The acquired video data is transmitted to the server via the network. Within the server, computational means abstract the visual information in a way that does not allow for individual identification. Specifically, software such as OpenCV and TensorFlow is used to abstract facial features and remove personally identifiable information from the data.

[0447] The converted visual information is processed by a generation device. This processed information is presented by display devices such as smart glasses or display devices, ensuring privacy in public spaces. Through the connected device, users can view the abstracted visual information and effectively understand the situation in public places while protecting their privacy.

[0448] A concrete example is a security monitoring system in a commercial facility. Cameras installed within the facility acquire visual information, and a server processes this in real time to abstract its features, enabling monitoring of the facility while thoroughly protecting privacy. Furthermore, prompts such as "Design an AI algorithm that recognizes and anonymizes faces from video footage taken in the city" can be used to help effectively build generative AI models.

[0449] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0450] Step 1:

[0451] The device acquires visual information in real time using a camera. In this case, the input is the surrounding video, and the output is raw video data. The camera sensor captures the video and stores it as high-resolution image data.

[0452] Step 2:

[0453] The terminal transmits raw video data to the server over the network. The input is the raw video data acquired by the terminal, and the output is the video data received by the server. A communication protocol is used to ensure that the data is transmitted accurately.

[0454] Step 3:

[0455] The server starts a computing system to process the received video data. The input is the video data received by the server, and the output is abstracted video data. Here, face recognition is performed using OpenCV, and features are abstracted using TensorFlow. Features necessary for identifying individuals are abstracted industrially, preventing individual identification.

[0456] Step 4:

[0457] The server generates abstracted images based on a generative AI model. The input is abstracted video data, and the output is the generated visual information. The model transforms the data into avatars or mosaics, representing it in a way that protects privacy.

[0458] Step 5:

[0459] The server sends the generated visual information back to the terminal. The input is the generated visual information, and the output is the abstracted image displayed on the user's terminal. The transmission protocol ensures that the data is transferred in the correct format.

[0460] Step 6:

[0461] The user perceives abstracted visual information through the device's display mechanism. The input is the abstracted image displayed on the device, while the output is the visual information the user perceives visually. The display device uses a high-resolution screen to visually represent detailed movements and situations.

[0462] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0463] This invention provides a system combining sensor means, computation means, generation means, display means, and an emotion engine, designed to protect user privacy and recognize emotions. This system acquires images of the environment using a camera installed on a terminal. The acquired image data is transmitted in real time to a server's computation means and processed by an emotion engine, including AI.

[0464] The server's computational means abstracts the people in the video and converts them into unidentifiable shapes, while simultaneously recognizing the user's emotions using an emotion engine. The emotion engine analyzes subtle facial movements and identifies the emotional state from the user's expressions and actions. Based on this emotional information, the generation means can adjust the video to suit the user.

[0465] As a concrete example, consider a security system installed at a company's reception area. The reception camera, which acts as the terminal, captures the visitor's image, but for security reasons, the video is always abstracted and transformed into a form that makes it impossible to identify an individual. Furthermore, an emotion engine analyzes the visitor's facial expressions and infers emotions such as interest or anxiety. If the system detects anxiety in the visitor, it notifies the reception staff, enabling a quick response. In this way, this system can understand the visitor's emotions while protecting individual privacy and improving the user experience.

[0466] Users can not only view real-time, privacy-protected video through their device's monitor, but also obtain additional information through emotion recognition. This system is effective in enhancing security and service quality in various environments such as commercial facilities, hospitals, and educational institutions.

[0467] The following describes the processing flow.

[0468] Step 1:

[0469] The device activates its camera, which is a sensor, and captures the surrounding image in real time. The camera captures the image data frame by frame and stores it in an internal buffer.

[0470] Step 2:

[0471] The terminal compresses the acquired video data and sends it to the server over the network. The data is transmitted using an appropriate protocol (e.g., RTSP) to ensure low latency.

[0472] Step 3:

[0473] The server decodes the video data received from the terminal and prepares it for processing as individual frames. The decoded frames are then formatted appropriately for AI processing and sentiment analysis.

[0474] Step 4:

[0475] The server's processing system analyzes video frames using image processing algorithms, abstracting them into shapes that make it impossible to identify individuals. This process is for privacy protection, ensuring that individuals cannot be identified.

[0476] Step 5:

[0477] The server further analyzes the user's emotions from abstracted video frames using an emotion engine. It detects subtle facial movements and estimates emotions such as joy, anger, sadness, and happiness based on them.

[0478] Step 6:

[0479] The server generation method incorporates the results of emotion analysis into abstracted images and adjusts the image content as needed. This adjustment aims for optimal display in specific emotional states.

[0480] Step 7:

[0481] The generated video is compressed again and sent from the server to the terminal. The terminal decodes it and displays it on the monitor. The user can view this video and utilize additional information obtained through emotion recognition while monitoring a privacy-protected environment.

[0482] Step 8:

[0483] Users monitor the video displayed on the screen in real time, checking for any abnormalities and emotional states as needed. If specific emotional actions are required, users can respond quickly.

[0484] (Example 2)

[0485] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0486] Conventional video processing systems have struggled to simultaneously protect privacy and recognize user emotions. While abstracting video footage from surveillance cameras and sensors is necessary to protect individual privacy, this process inevitably leads to a decrease in the accuracy of emotion recognition. Furthermore, performing these processes in real time requires significant system processing power.

[0487] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0488] In this invention, the server includes detection means for acquiring video, computation means for processing and abstracting the video in real time, and analysis means for performing emotion recognition. This makes it possible to accurately recognize the user's emotions in real time while protecting the user's privacy.

[0489] "Detection means" refers to devices or sensors used to acquire video data from the environment.

[0490] "Computation means" refers to a device or process for processing acquired video data and performing abstraction as necessary.

[0491] "Production means" refers to devices or methods for generating processed and abstracted video data for subsequent processing or display.

[0492] "Presentation means" refers to a device or interface for visually providing generated video data to the user.

[0493] "Analysis means" refers to a device or method that recognizes a person's emotions based on acquired and processed video data and outputs the analysis results.

[0494] This invention provides a system for emotion recognition while protecting user privacy. The system primarily functions by acquiring video footage from a camera installed on a terminal and processing that data on a server. The following describes a specific implementation of this system.

[0495] Terminal role:

[0496] The device is equipped with a high-sensitivity camera for acquiring video. This camera can capture the environment and user's movements in real time. The device also has a communication module for transferring the acquired video data to a server, and the data is transmitted using a secure protocol.

[0497] Server role:

[0498] The server is equipped with high-performance computing devices to process the received video data. These devices utilize AI models as computational tools to perform data abstraction and emotion analysis. Video abstraction is essential from a privacy perspective, specifically transforming the images into shapes that make it impossible to identify individuals. Furthermore, the emotion engine has the ability to identify emotional states from the user's facial expressions and movements.

[0499] User roles:

[0500] Users can view privacy-protected video and analyzed sentiment information through their device's monitor. This enables users to respond appropriately and provide high-quality service to visitors.

[0501] Specific example:

[0502] For example, this system could be implemented in a company's reception area. In this case, the visitor's image would always be abstracted, making it impossible to identify an individual. On the other hand, the emotion engine would analyze the visitor's facial expressions and determine emotions such as interest or anxiety. If the system determines that a visitor is showing anxiety, the reception staff could respond to the visitor quickly.

[0503] Example of a prompt:

[0504] "Please describe in detail how you receive visitor facial expression data and analyze emotions using your emotion engine. Also, please describe how the system uses this analysis to improve the user experience while protecting privacy."

[0505] This system enables the provision of high-quality services while maintaining privacy in commercial facilities, medical institutions, educational institutions, and other settings.

[0506] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0507] Step 1:

[0508] The device uses a camera to acquire video footage of the environment in real time. At this point, the input is visual information of the surroundings. The device converts this video data into a digital format and sends it to the server in real time. Specifically, the camera continuously captures frames and transfers them to the server via a secure communication protocol.

[0509] Step 2:

[0510] The server inputs the received video data into a computing device. Here, initial data processing is performed to analyze the video data and extract important information. This process involves adjusting the video resolution and performing pre-processing such as noise reduction. The output is processed, clear video data.

[0511] Step 3:

[0512] The server's computing power abstracts the processed video data. Specifically, it uses machine learning algorithms to detect people in the video and transforms their features into shapes that are unidentifiable. The input for this step is preprocessed video data, and the output is abstracted video. The system achieves privacy protection at this stage.

[0513] Step 4:

[0514] The server inputs abstracted video into its emotion engine and performs detailed emotion analysis. The analysis method uses a machine learning model to infer emotions based on subtle facial changes in the video. The input for this step is abstracted video, and the output is identified emotion information. Specific operations include facial feature point extraction and emotional state classification.

[0515] Step 5:

[0516] The server inputs the extracted emotional information into a production mechanism that generates content for the user. This mechanism adjusts the images and information as needed based on the emotional information and presents it in a way that is useful to the user. The output of this step is information and images customized for the user. Specifically, this involves adjusting the screen display and creating notifications according to the emotional state.

[0517] Step 6:

[0518] Users view protected video and emotional information in real time through the presentation tools. They use this information to respond appropriately to visitors. The input in this step is customized information, and the output is the user's actions. Specific actions include the user providing situation-appropriate customer service and responses.

[0519] (Application Example 2)

[0520] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0521] In conventional stores, it has been difficult to analyze customers' facial expressions and behavior in real time and flexibly adjust customer service based on that analysis. Furthermore, it has been challenging to identify customer emotions while protecting privacy during video processing. This invention aims to solve these problems and realize efficient and privacy-conscious customer service that responds to customer emotions.

[0522] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0523] In this invention, the server includes detection means for acquiring video, computation means for processing the video acquired by the detection means, and creation means for generating an abstracted video by the computation means. This makes it possible to analyze customer emotions based on video that has been transformed into shapes that do not allow for the identification of individuals, and to present that information in real time.

[0524] "Detection means for acquiring images" refers to a sensor device for acquiring visual information of the physical environment.

[0525] "Computational means for processing the video acquired by the detection means" refers to a computer system for analyzing the acquired video data and extracting specific information.

[0526] "The calculation means that generates an abstracted image" refers to a device that provides a process for processing an image from data analyzed by the calculation means in a way that prevents the identification of individuals.

[0527] "Presentation means for displaying the image generated by the creation means" refers to a device having a display function that makes the abstracted image viewable by the user or system.

[0528] "An emotion analysis tool for analyzing a customer's emotional state" refers to an algorithm or system that identifies and determines an emotion from the facial expressions and actions of a person in a video.

[0529] "Information presentation means for displaying emotional information analyzed by the emotional analysis means" refers to a device or system for providing information obtained through emotional analysis to a user visually or by other means.

[0530] The system of this invention is built to improve the customer experience in physical stores. The system includes the following hardware and software configuration.

[0531] Hardware configuration:

[0532] Detection devices capable of acquiring information at high resolution (e.g., cameras)

[0533] Visual presentation devices equipped with video output displays (e.g., smart glasses)

[0534] Software configuration:

[0535] A computational program for image processing (e.g., OpenCV library)

[0536] An emotion analysis processor that analyzes emotions from a customer's facial expressions (e.g., Microsoft Azure Emotion API)

[0537] The server first acquires images of customers using detection devices installed in the store. Then, it processes the acquired video data using computational means, abstracting it in a way that protects individual privacy while preventing identification. Next, an emotion analysis means analyzes the customer's emotional state from the abstracted video, and the results are displayed in real time on a visual display device worn by staff using an information display means. This allows staff to instantly understand the customer's interests and anxieties, enabling them to take more appropriate action.

[0538] For example, when a new product is introduced in a store, if a customer is looking at the new product with interest, the visual display device will show an analysis result indicating "showing interest." This allows staff to proactively initiate customer service, explaining and suggesting the product.

[0539] Examples of prompts include, "Analyze the facial expressions of customers in the store and analyze their emotions, such as interest and anxiety, in real time," and "Based on the changes in customer emotions, suggest appropriate countermeasures to the service staff."

[0540] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0541] Step 1:

[0542] The server acquires video from detection devices installed in physical stores. The input is real-time camera footage, and the output is saved as video data. The server immediately prepares to place this video data into a processing queue.

[0543] Step 2:

[0544] The server analyzes the acquired video using computational means. The input is the video data acquired in step 1, and the output is abstracted video data. The server uses an image processing library (e.g., OpenCV) to abstract people in the video and convert them into shapes that do not allow for individual identification.

[0545] Step 3:

[0546] The server feeds abstracted video data into an emotion analysis system to analyze the customer's emotional state. The input is the abstracted data generated in step 2, and the output is the analyzed emotional information. A generative AI model is used to identify emotions from the abstracted facial features and extract those states as numerical data or labels.

[0547] Step 4:

[0548] The server transmits the analyzed emotion information to a visual display device, and the user provides customer service based on the information displayed on the device. The input is the emotion information obtained in step 3, and the output is the emotion label or numerical information displayed on the visual display device. Based on this, the user considers appropriate suggestions and customer service strategies for the customer.

[0549] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0550] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0551] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0552] [Fourth Embodiment]

[0553] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0554] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0555] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0556] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0557] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0558] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0559] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0560] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0561] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0562] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0563] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0564] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0565] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0566] As an embodiment for carrying out the present invention, a surveillance system for the purpose of protecting privacy will be described. This system acquires images of the environment using a camera, which is a sensor means installed at a terminal, and processes the images in real time using a computing means on a server. The computing means uses an advanced image processing algorithm to abstract people in the images and convert them into shapes that make it impossible to identify individuals. As a result of this conversion process, images that protect privacy are output by a generation means.

[0567] As a concrete example, let's describe a surveillance camera system installed in a public facility. The camera, acting as the terminal, continuously acquires video footage from the location. This data is transmitted to a server via a network. The server's computing unit processes the incoming video data in real time, using AI technology to abstract the appearance of people. The generation unit generates the processed video based on this abstraction, and the display unit projects the abstracted video onto a monitor connected to the terminal.

[0568] Users can view privacy-protected video footage through their device's monitor. While the video allows for the identification of human movement and location, individual features are abstracted, making it impossible to identify specific individuals. Furthermore, this system is adaptable to diverse environments and is effective for surveillance in commercial facilities, hospitals, and educational institutions. Additionally, since the video is processed in real time, it can handle situations requiring immediate monitoring.

[0569] The following describes the processing flow.

[0570] Step 1:

[0571] The device activates its camera sensor to acquire video from the environment, capturing surrounding visual data in real time. The camera processes the video data frame by frame and stores it in an internal buffer.

[0572] Step 2:

[0573] The terminal transmits the acquired video data to the server via the network. Transmission takes place in real time, and the video data is encoded in an appropriate compression format (e.g., H.264) to ensure communication efficiency.

[0574] Step 3:

[0575] The server decodes the video data received from the terminal and prepares it for processing as individual frames. The decoded data is then converted into a format suitable for the image processing algorithm.

[0576] Step 4:

[0577] The server processes the video data frame by frame using an AI model. The AI ​​recognizes the shape and movement of people and converts them into unidentifiable silhouettes or abstract shapes.

[0578] Step 5:

[0579] The server re-encodes the abstracted video frames processed by the AI, generates them in real time using a generation mechanism, and sends them to the terminal. The network is optimized to ensure that the transmitted data arrives with low latency.

[0580] Step 6:

[0581] The terminal decodes abstracted video transmitted from the server and displays it in real time on a connected display device. This allows the user to monitor a privacy-protected environment.

[0582] Step 7:

[0583] Users can monitor the video displayed on their device to check for any unusual behavior or situations. They can also replay saved footage for further review if necessary. This allows users to use the camera system with peace of mind.

[0584] (Example 1)

[0585] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0586] In recent years, the need for privacy protection in surveillance systems has increased. However, conventional surveillance systems record individuals' appearances directly, which can lead to privacy violations. In addition, while surveillance data needs to be processed in real time, problems with processing delays and accuracy arise. To solve these problems, it is necessary to build a system that records data in a way that makes it impossible to identify individuals and processes it quickly.

[0587] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0588] In this invention, the server includes a measuring device that acquires video, an information processing device that processes the video acquired by the measuring device in real time and extracts characteristic points of a person, and a reconstruction device that generates an abstracted video from the information processing device. This makes it possible to efficiently process the video in real time while converting the appearance of an individual in the video into an unidentifiable shape.

[0589] A "measuring device" is a device used to acquire images of the environment and has the role of transmitting the image information to other devices.

[0590] An "information processing device" is a device that processes video data acquired from a measuring device in real time and extracts characteristic points of people in the video.

[0591] A "reconstruction device" is a device that generates abstract images that do not allow for the identification of individuals, based on feature points extracted by an information processing device.

[0592] A "monitoring system" refers to an entire system that combines measuring devices, information processing devices, and reconfiguration devices to monitor the environment while protecting privacy.

[0593] A "generative AI model" is a model that utilizes artificial intelligence technology used in information processing devices to achieve abstraction of video data.

[0594] This invention relates to a system for monitoring the environment while protecting privacy. The system mainly consists of a measuring device, an information processing device, and a reconstruction device, which are combined to implement the system.

[0595] The terminal unit's measuring device includes a camera for acquiring high-resolution images of the environment, and acquires video data in real time. This data is transmitted to a server via the network.

[0596] The server processes the received video data in real time using an information processing device. During this process, it utilizes a generative AI model to extract characteristic features of individuals and transforms these features into shapes that cannot identify individuals. The information processing device has advanced parallel processing capabilities and is designed to minimize processing delays.

[0597] The reconstruction device generates a new image based on abstracted feature points. This generated image is output to a monitor via a display device connected to the terminal, in a privacy-protected manner. The user can monitor through this privacy-protected image and take action as needed.

[0598] For example, in commercial facilities, the terminal's camera can capture the flow of customers, which can be used to understand congestion levels and consider safety measures. This system allows users, who are managers of commercial facilities, to understand the movements within the facility while ensuring the privacy of customers.

[0599] An example of a prompt message is: "Develop a program that abstracts individual characteristics for privacy protection and displays the video in real time."

[0600] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0601] Step 1:

[0602] The terminal's measuring device acquires video of the environment. It captures high-resolution video data in real time using an optical sensor as input. Specifically, it automatically adjusts exposure and focus according to ambient lighting conditions to improve data acquisition accuracy. The output is a stream of acquired video data, ready to be sent to the next processing step.

[0603] Step 2:

[0604] The terminal transmits the acquired video data to the server via the network. The input is the video data acquired in step 1. Specifically, the data is compressed and packetized, and stable data transfer is achieved through the communication protocol. As output, the data is transmitted to the server within the expected bandwidth.

[0605] Step 3:

[0606] The server's information processing unit receives the acquired video data in real time and begins processing it. The input is video data transmitted over the network. Using a generative AI model, it performs data calculations to extract feature points of people from the video, and based on this information, generates abstract data that makes it impossible to identify individuals. The output is video data abstracted based on the feature points.

[0607] Step 4:

[0608] The server's reconstruction device generates new video based on the abstracted video data. The input is the abstracted data obtained in step 3. Specifically, it applies mosaic and blurring processes to reconstruct the video in a state where people cannot be identified. The output is privacy-protected video data that can be used by display devices.

[0609] Step 5:

[0610] The terminal displays privacy-protected video footage sent from the server on its monitor. The input is video data generated by a reconstruction device. The specific operation involves adjusting the resolution and color correction to display the acquired video in a clear manner. The output appears on the monitor as an abstracted image that the user can view.

[0611] Step 6:

[0612] The user monitors the privacy-protected video displayed on the monitor. The input is the video displayed in step 5. Specific actions include rewinding the video as needed and focusing on specific areas to analyze movement patterns. The output is the decision based on the monitoring and the plan for implementing safety measures.

[0613] (Application Example 1)

[0614] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0615] With the widespread use of devices that acquire visual information, protecting privacy in public spaces has become increasingly important. However, conventional methods often display the visual characteristics of others in an identifiable way, making it difficult to completely protect individual privacy. Therefore, there is a need for a system that converts visual information in public spaces into a form that does not allow for personal identification, thereby enabling safe and secure information verification.

[0616] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0617] In this invention, the server includes detection means for extracting visual information, calculation means for transforming the visual information acquired by the detection means, and generation means for displaying the visual information generated by the calculation means. This makes it possible to abstract the visual information and confirm the situation in public spaces while preventing the identification of individuals.

[0618] "Visual information" refers to image and video data acquired by cameras and other light detection devices.

[0619] "Detection means" refers to sensors or devices used to collect visual information.

[0620] "Computational means" refers to algorithms and devices used to process acquired visual information and perform transformations and analyses as needed.

[0621] "Generating means" refers to devices or processes for creating a new form from visual information processed by computation means.

[0622] "Display means" refers to screens or display devices used to visually convey generated visual information to the user.

[0623] A "connection device" is an interface or network device used to exchange information with other devices or systems.

[0624] This invention provides a system that can effectively acquire visual information and confirm a situation while preventing the identification of personal information. To carry out the invention, a system including the following components is required.

[0625] The server uses cameras and other light detection devices as detection means to extract visual information. These detection means are placed in public spaces and acquire video in real time. The acquired video data is transmitted to the server via the network. Within the server, computational means abstract the visual information in a way that does not allow for individual identification. Specifically, software such as OpenCV and TensorFlow is used to abstract facial features and remove personally identifiable information from the data.

[0626] The converted visual information is processed by a generation device. This processed information is presented by display devices such as smart glasses or display devices, ensuring privacy in public spaces. Through the connected device, users can view the abstracted visual information and effectively understand the situation in public places while protecting their privacy.

[0627] A concrete example is a security monitoring system in a commercial facility. Cameras installed within the facility acquire visual information, and a server processes this in real time to abstract its features, enabling monitoring of the facility while thoroughly protecting privacy. Furthermore, prompts such as "Design an AI algorithm that recognizes and anonymizes faces from video footage taken in the city" can be used to help effectively build generative AI models.

[0628] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0629] Step 1:

[0630] The device acquires visual information in real time using a camera. In this case, the input is the surrounding video, and the output is raw video data. The camera sensor captures the video and stores it as high-resolution image data.

[0631] Step 2:

[0632] The terminal transmits raw video data to the server over the network. The input is the raw video data acquired by the terminal, and the output is the video data received by the server. A communication protocol is used to ensure that the data is transmitted accurately.

[0633] Step 3:

[0634] The server starts a computing system to process the received video data. The input is the video data received by the server, and the output is abstracted video data. Here, face recognition is performed using OpenCV, and features are abstracted using TensorFlow. Features necessary for identifying individuals are abstracted industrially, preventing individual identification.

[0635] Step 4:

[0636] The server generates abstracted images based on a generative AI model. The input is abstracted video data, and the output is the generated visual information. The model transforms the data into avatars or mosaics, representing it in a way that protects privacy.

[0637] Step 5:

[0638] The server sends the generated visual information back to the terminal. The input is the generated visual information, and the output is the abstracted image displayed on the user's terminal. The transmission protocol ensures that the data is transferred in the correct format.

[0639] Step 6:

[0640] The user perceives abstracted visual information through the device's display mechanism. The input is the abstracted image displayed on the device, while the output is the visual information the user perceives visually. The display device uses a high-resolution screen to visually represent detailed movements and situations.

[0641] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0642] This invention provides a system combining sensor means, computation means, generation means, display means, and an emotion engine, designed to protect user privacy and recognize emotions. This system acquires images of the environment using a camera installed on a terminal. The acquired image data is transmitted in real time to a server's computation means and processed by an emotion engine, including AI.

[0643] The server's computational means abstracts the people in the video and converts them into unidentifiable shapes, while simultaneously recognizing the user's emotions using an emotion engine. The emotion engine analyzes subtle facial movements and identifies the emotional state from the user's expressions and actions. Based on this emotional information, the generation means can adjust the video to suit the user.

[0644] As a concrete example, consider a security system installed at a company's reception area. The reception camera, which acts as the terminal, captures the visitor's image, but for security reasons, the video is always abstracted and transformed into a form that makes it impossible to identify an individual. Furthermore, an emotion engine analyzes the visitor's facial expressions and infers emotions such as interest or anxiety. If the system detects anxiety in the visitor, it notifies the reception staff, enabling a quick response. In this way, this system can understand the visitor's emotions while protecting individual privacy and improving the user experience.

[0645] Users can not only view real-time, privacy-protected video through their device's monitor, but also obtain additional information through emotion recognition. This system is effective in enhancing security and service quality in various environments such as commercial facilities, hospitals, and educational institutions.

[0646] The following describes the processing flow.

[0647] Step 1:

[0648] The device activates its camera, which is a sensor, and captures the surrounding image in real time. The camera captures the image data frame by frame and stores it in an internal buffer.

[0649] Step 2:

[0650] The terminal compresses the acquired video data and sends it to the server over the network. The data is transmitted using an appropriate protocol (e.g., RTSP) to ensure low latency.

[0651] Step 3:

[0652] The server decodes the video data received from the terminal and prepares it for processing as individual frames. The decoded frames are then formatted appropriately for AI processing and sentiment analysis.

[0653] Step 4:

[0654] The server's processing system analyzes video frames using image processing algorithms, abstracting them into shapes that make it impossible to identify individuals. This process is for privacy protection, ensuring that individuals cannot be identified.

[0655] Step 5:

[0656] The server further analyzes the user's emotions from abstracted video frames using an emotion engine. It detects subtle facial movements and estimates emotions such as joy, anger, sadness, and happiness based on them.

[0657] Step 6:

[0658] The server generation method incorporates the results of emotion analysis into abstracted images and adjusts the image content as needed. This adjustment aims for optimal display in specific emotional states.

[0659] Step 7:

[0660] The generated video is compressed again and sent from the server to the terminal. The terminal decodes it and displays it on the monitor. The user can view this video and utilize additional information obtained through emotion recognition while monitoring a privacy-protected environment.

[0661] Step 8:

[0662] Users monitor the video displayed on the screen in real time, checking for any abnormalities and emotional states as needed. If specific emotional actions are required, users can respond quickly.

[0663] (Example 2)

[0664] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0665] Conventional video processing systems have struggled to simultaneously protect privacy and recognize user emotions. While abstracting video footage from surveillance cameras and sensors is necessary to protect individual privacy, this process inevitably leads to a decrease in the accuracy of emotion recognition. Furthermore, performing these processes in real time requires significant system processing power.

[0666] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0667] In this invention, the server includes detection means for acquiring video, computation means for processing and abstracting the video in real time, and analysis means for performing emotion recognition. This makes it possible to accurately recognize the user's emotions in real time while protecting the user's privacy.

[0668] "Detection means" refers to devices or sensors used to acquire video data from the environment.

[0669] "Computation means" refers to a device or process for processing acquired video data and performing abstraction as necessary.

[0670] "Production means" refers to devices or methods for generating processed and abstracted video data for subsequent processing or display.

[0671] "Presentation means" refers to a device or interface for visually providing generated video data to the user.

[0672] "Analysis means" refers to a device or method that recognizes a person's emotions based on acquired and processed video data and outputs the analysis results.

[0673] This invention provides a system for emotion recognition while protecting user privacy. The system primarily functions by acquiring video footage from a camera installed on a terminal and processing that data on a server. The following describes a specific implementation of this system.

[0674] Terminal role:

[0675] The device is equipped with a high-sensitivity camera for acquiring video. This camera can capture the environment and user's movements in real time. The device also has a communication module for transferring the acquired video data to a server, and the data is transmitted using a secure protocol.

[0676] Server role:

[0677] The server is equipped with high-performance computing devices to process the received video data. These devices utilize AI models as computational tools to perform data abstraction and emotion analysis. Video abstraction is essential from a privacy perspective, specifically transforming the images into shapes that make it impossible to identify individuals. Furthermore, the emotion engine has the ability to identify emotional states from the user's facial expressions and movements.

[0678] User roles:

[0679] Users can view privacy-protected video and analyzed sentiment information through their device's monitor. This enables users to respond appropriately and provide high-quality service to visitors.

[0680] Specific example:

[0681] For example, this system could be implemented in a company's reception area. In this case, the visitor's image would always be abstracted, making it impossible to identify an individual. On the other hand, the emotion engine would analyze the visitor's facial expressions and determine emotions such as interest or anxiety. If the system determines that a visitor is showing anxiety, the reception staff could respond to the visitor quickly.

[0682] Example of a prompt:

[0683] "Please describe in detail how you receive visitor facial expression data and analyze emotions using your emotion engine. Also, please describe how the system uses this analysis to improve the user experience while protecting privacy."

[0684] This system enables the provision of high-quality services while maintaining privacy in commercial facilities, medical institutions, educational institutions, and other settings.

[0685] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0686] Step 1:

[0687] The device uses a camera to acquire video footage of the environment in real time. At this point, the input is visual information of the surroundings. The device converts this video data into a digital format and sends it to the server in real time. Specifically, the camera continuously captures frames and transfers them to the server via a secure communication protocol.

[0688] Step 2:

[0689] The server inputs the received video data into a computing device. Here, initial data processing is performed to analyze the video data and extract important information. This process involves adjusting the video resolution and performing pre-processing such as noise reduction. The output is processed, clear video data.

[0690] Step 3:

[0691] The server's computing power abstracts the processed video data. Specifically, it uses machine learning algorithms to detect people in the video and transforms their features into shapes that are unidentifiable. The input for this step is preprocessed video data, and the output is abstracted video. The system achieves privacy protection at this stage.

[0692] Step 4:

[0693] The server inputs abstracted video into its emotion engine and performs detailed emotion analysis. The analysis method uses a machine learning model to infer emotions based on subtle facial changes in the video. The input for this step is abstracted video, and the output is identified emotion information. Specific operations include facial feature point extraction and emotional state classification.

[0694] Step 5:

[0695] The server inputs the extracted emotional information into a production mechanism that generates content for the user. This mechanism adjusts the images and information as needed based on the emotional information and presents it in a way that is useful to the user. The output of this step is information and images customized for the user. Specifically, this involves adjusting the screen display and creating notifications according to the emotional state.

[0696] Step 6:

[0697] Users view protected video and emotional information in real time through the presentation tools. They use this information to respond appropriately to visitors. The input in this step is customized information, and the output is the user's actions. Specific actions include the user providing situation-appropriate customer service and responses.

[0698] (Application Example 2)

[0699] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0700] In conventional stores, it has been difficult to analyze customers' facial expressions and behavior in real time and flexibly adjust customer service based on that analysis. Furthermore, it has been challenging to identify customer emotions while protecting privacy during video processing. This invention aims to solve these problems and realize efficient and privacy-conscious customer service that responds to customer emotions.

[0701] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0702] In this invention, the server includes detection means for acquiring video, computation means for processing the video acquired by the detection means, and creation means for generating an abstracted video by the computation means. This makes it possible to analyze customer emotions based on video that has been transformed into shapes that do not allow for the identification of individuals, and to present that information in real time.

[0703] "Detection means for acquiring images" refers to a sensor device for acquiring visual information of the physical environment.

[0704] "Computational means for processing the video acquired by the detection means" refers to a computer system for analyzing the acquired video data and extracting specific information.

[0705] "The calculation means that generates an abstracted image" refers to a device that provides a process for processing an image from data analyzed by the calculation means in a way that prevents the identification of individuals.

[0706] "Presentation means for displaying the image generated by the creation means" refers to a device having a display function that makes the abstracted image viewable by the user or system.

[0707] "An emotion analysis tool for analyzing a customer's emotional state" refers to an algorithm or system that identifies and determines an emotion from the facial expressions and actions of a person in a video.

[0708] "Information presentation means for displaying emotional information analyzed by the emotional analysis means" refers to a device or system for providing information obtained through emotional analysis to a user visually or by other means.

[0709] The system of this invention is built to improve the customer experience in physical stores. The system includes the following hardware and software configuration.

[0710] Hardware configuration:

[0711] Detection devices capable of acquiring information at high resolution (e.g., cameras)

[0712] Visual presentation devices equipped with video output displays (e.g., smart glasses)

[0713] Software configuration:

[0714] A computational program for image processing (e.g., OpenCV library)

[0715] An emotion analysis processor that analyzes emotions from a customer's facial expressions (e.g., Microsoft Azure Emotion API)

[0716] The server first acquires images of customers using detection devices installed in the store. Then, it processes the acquired video data using computational means, abstracting it in a way that protects individual privacy while preventing identification. Next, an emotion analysis means analyzes the customer's emotional state from the abstracted video, and the results are displayed in real time on a visual display device worn by staff using an information display means. This allows staff to instantly understand the customer's interests and anxieties, enabling them to take more appropriate action.

[0717] For example, when a new product is introduced in a store, if a customer is looking at the new product with interest, the visual display device will show an analysis result indicating "showing interest." This allows staff to proactively initiate customer service, explaining and suggesting the product.

[0718] Examples of prompts include, "Analyze the facial expressions of customers in the store and analyze their emotions, such as interest and anxiety, in real time," and "Based on the changes in customer emotions, suggest appropriate countermeasures to the service staff."

[0719] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0720] Step 1:

[0721] The server acquires video from detection devices installed in physical stores. The input is real-time camera footage, and the output is saved as video data. The server immediately prepares to place this video data into a processing queue.

[0722] Step 2:

[0723] The server analyzes the acquired video using computational means. The input is the video data acquired in step 1, and the output is abstracted video data. The server uses an image processing library (e.g., OpenCV) to abstract people in the video and convert them into shapes that do not allow for individual identification.

[0724] Step 3:

[0725] The server feeds abstracted video data into an emotion analysis system to analyze the customer's emotional state. The input is the abstracted data generated in step 2, and the output is the analyzed emotional information. A generative AI model is used to identify emotions from the abstracted facial features and extract those states as numerical data or labels.

[0726] Step 4:

[0727] The server transmits the analyzed emotion information to a visual display device, and the user provides customer service based on the information displayed on the device. The input is the emotion information obtained in step 3, and the output is the emotion label or numerical information displayed on the visual display device. Based on this, the user considers appropriate suggestions and customer service strategies for the customer.

[0728] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0729] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0730] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0731] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0732] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0733] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0734] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0735] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0736] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0737] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0738] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0739] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0740] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0741] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0742] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0743] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0744] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0745] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0746] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0747] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0748] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0749] The following is further disclosed regarding the embodiments described above.

[0750] (Claim 1)

[0751] A sensor means for acquiring video,

[0752] A calculation means for processing the video acquired by the sensor means,

[0753] The calculation means generates an abstracted image,

[0754] A display means for displaying the video generated by the generation means,

[0755] A system that includes this.

[0756] (Claim 2)

[0757] The system according to claim 1, wherein the generating means converts the characteristics of a person into a shape that makes identification impossible.

[0758] (Claim 3)

[0759] The system according to claim 1, wherein the calculation means performs processing to abstract the video in real time.

[0760] "Example 1"

[0761] (Claim 1)

[0762] A measuring device that acquires images,

[0763] An information processing device that processes video acquired by the measuring device in real time and extracts characteristic points of a person,

[0764] The information processing device is a reconstruction device that generates an abstracted image,

[0765] A display device that displays the image generated by the reconstruction device,

[0766] A monitoring system including this.

[0767] (Claim 2)

[0768] The surveillance system according to claim 1, wherein the reconstruction device converts the appearance of a person into a shape that does not allow for personal identification through abstraction processing.

[0769] (Claim 3)

[0770] The surveillance system according to claim 1, wherein the information processing device performs a process of abstracting video using a generated AI model.

[0771] "Application Example 1"

[0772] (Claim 1)

[0773] A detection means for acquiring visual information,

[0774] A computing means for processing visual information acquired by the detection means,

[0775] The calculation means is a generation means that generates abstracted visual information,

[0776] A display means for displaying the visual information generated by the generation means,

[0777] A system including a connecting device that abstracts the visual information displayed on the display means to prevent the identification of people.

[0778] (Claim 2)

[0779] The system according to claim 1, wherein the generating means converts visual features into an unidentifiable shape.

[0780] (Claim 3)

[0781] The system according to claim 1, wherein the calculation means performs a process to abstract visual information in real time and allows the status of operation to be confirmed.

[0782] "Example 2 of combining an emotion engine"

[0783] (Claim 1)

[0784] A detection means for acquiring video,

[0785] A computing means for processing the video acquired by the detection means,

[0786] The calculation means is a production means that generates an abstracted image,

[0787] A presentation means for displaying the image generated by the production means,

[0788] An analytical means for performing emotion recognition,

[0789] A system that includes this.

[0790] (Claim 2)

[0791] The system according to claim 1, wherein the production means performs a process that converts the characteristics of a person into a shape that cannot be identified.

[0792] (Claim 3)

[0793] The system according to claim 1, wherein the calculation means simultaneously performs a process to abstract the video in real time and a process to recognize the user's emotions.

[0794] "Application example 2 when combining with an emotional engine"

[0795] (Claim 1)

[0796] A detection means for acquiring video,

[0797] A computing means for processing the video acquired by the detection means,

[0798] The calculation means is a creation means that generates an abstracted image,

[0799] A presentation means for displaying the video generated by the creation means,

[0800] A means of analyzing the emotional state of customers,

[0801] An information presentation means that displays emotional information analyzed by the emotional analysis means,

[0802] A system that includes this.

[0803] (Claim 2)

[0804] The system according to claim 1, wherein the creation means converts the characteristics of a person into a shape that makes identification impossible.

[0805] (Claim 3)

[0806] The system according to claim 1, wherein the calculation means performs a process to abstract the video in real time, and the information presentation means presents the customer's emotional state in real time. [Explanation of symbols]

[0807] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A sensor means for acquiring video, A calculation means for processing the video acquired by the sensor means, The calculation means generates an abstracted image, A display means for displaying the video generated by the generation means, A system that includes this.

2. The system according to claim 1, wherein the generation means converts the characteristics of a person into a shape that makes identification impossible.

3. The system according to claim 1, wherein the calculation means performs processing to abstract the video in real time.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A