Information processing device, information processing method, and recording medium
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-01-27
- Publication Date
- 2026-08-13
Smart Images

Figure JP2026002583_13082026_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and recording medium
[0001] This disclosure relates to an information processing device, an information processing method, a program, and a recording medium.
[0002] A technology related to this disclosure is disclosed in Patent Document 1. Patent Document 1 discloses a method of identifying the content of what people are saying in an image by analyzing security camera images using lip-reading, and displaying the content of what each person is saying in a speech bubble on the image.
[0003] Japanese Patent Publication No. 2010-283637
[0004] There are operations that involve monitoring images captured by cameras. One example of the purpose of this disclosure is to advance the technology related to such operations.
[0005] According to one aspect of this disclosure, an information processing device is provided, comprising: acquisition means for acquiring a target image; statement information generation means for generating statement information relating to statements made by a person included in the target image; state information generation means for generating state information relating to the state of a person included in the target image; detection means for detecting the occurrence of a problem event based on at least one of the statement information and the state information; and output means for outputting the target image and the results of the detection.
[0006] Furthermore, according to one aspect of this disclosure, an information processing method is provided in which one or more computers acquire a target image, generate speech information relating to the speech of a person included in the target image, generate state information relating to the state of a person included in the target image, detect the occurrence of a problem event based on at least one of the speech information and the state information, and output the target image and the result of the detection.
[0007] Furthermore, according to one aspect of this disclosure, a program is provided that causes a computer to function as: an acquisition means for acquiring a target image; a speech information generation means for generating speech information relating to speeches made by a person included in the target image; a state information generation means for generating state information relating to the state of a person included in the target image; a detection means for detecting the occurrence of a problem event based on at least one of the speech information and the state information; and an output means for outputting the target image and the results of the detection.
[0008] One example of this disclosure suggests that technologies related to operations for monitoring images captured by cameras can be developed.
[0009] Figure 1 is a diagram showing an example of a functional block diagram of an information processing device. Figure 2 is a flowchart showing an example of the processing flow of an information processing device. Figure 3 is a diagram showing an example of the hardware configuration of an information processing device. Figure 4 is a diagram schematically showing an example of the information processed by an information processing device. Figure 5 is a diagram showing an example of the screen output by an information processing device. Figure 6 is a flowchart showing another example of the processing flow of an information processing device.
[0010] The embodiments of this disclosure will be described below with reference to the drawings. In this disclosure, the drawings are associated with one or more embodiments. In all drawings, similar components are denoted by the same reference numerals, and their descriptions are omitted where appropriate.
[0011] <<First Embodiment>> Figure 1 is a functional block diagram showing an overview of the information processing device 10. Figure 2 is a flowchart showing an example of the processing flow executed by the information processing device 10.
[0012] As shown in Figure 1, the information processing device 10 includes an acquisition unit 11, a speech information generation unit 12, a state information generation unit 13, a detection unit 14, and an output unit 15. These functional units execute the processes shown in the flowchart of Figure 2.
[0013] In S10, the acquisition unit 11 acquires the target image. In S11, the speech information generation unit 12 generates speech information regarding the speech of a person included in the target image. In S12, the state information generation unit 13 generates state information regarding the state of a person included in the target image. In S13, the detection unit 14 detects the occurrence of a problem event based on at least one of the speech information and the state information. In S14, the output unit 15 outputs the detection result.
[0014] Note that the processing order of S11 and S12 is not limited to this example. For example, S12 may be performed before S11, or S11 and S12 may be performed in parallel.
[0015] By the way, there is an operation to monitor images captured by cameras. For example, in a monitoring center, images captured by cameras are displayed on a monitor. Then, a monitor monitors those images. Note that sometimes only one image is displayed on the monitor, and sometimes multiple images are displayed. In the latter case, for example, multiple images taken by multiple cameras, each capturing a different location, are displayed side by side on the monitor.
[0016] In operations that monitor images captured by cameras, audio is often not output, and only the image is output. For example, in multi-screen monitoring, if audio is output from multiple locations, the sounds from each location will mix, making it difficult to clearly identify the content. Therefore, in multi-screen monitoring, audio is not output, and monitoring is performed using only images. In addition, some surveillance cameras do not have a microphone function at all. Therefore, even when monitoring by displaying a single image on a monitor, monitoring may be performed using only images, without outputting audio.
[0017] When monitoring is performed using only images and without audio output, the monitor must detect problems based solely on the image information. However, if audio information can be used in addition to the image information, problems can be detected with even greater accuracy. The information processing device 10 takes this into consideration and supports the operation of monitoring images captured by the camera.
[0018] As described above, the information processing device 10 detects the occurrence of a problem event based on "speech information related to the statements of a person included in the target image" and outputs the detection result. In other words, the information processing device 10 detects the occurrence of a problem event based on "audio information" that is unavailable to the monitor monitoring the target image and outputs the detection result. In this way, the information processing device 10 supports the monitor by detecting the occurrence of a problem event based on information unavailable to the monitor and providing the results to the monitor.
[0019] Furthermore, the information processing device 10 detects the occurrence of a problem event based on state information relating to the state of a person included in the target image, in addition to or instead of statement information relating to the person's statements included in the target image. The information processing device 10, which can utilize multiple characteristic pieces of information such as statement information and state information, can accurately detect the occurrence of a problem event.
[0020] Such an information processing device 10 makes it possible to develop technology related to the operation of monitoring images captured by a camera.
[0021] <<Second Embodiment>> <Overview> The information processing device 10 of the second embodiment is a concrete implementation of the configuration of the information processing device 10 of the first embodiment. It will be described in detail below.
[0022] <Hardware Configuration> First, an example of the hardware configuration of the information processing device 10 will be described. Each functional unit of the information processing device 10 is realized by any combination of hardware and software. Those skilled in the art will understand that there are various variations in the implementation method and the device. The software includes programs that are pre-installed at the time of shipment of the device, as well as programs downloaded from recording media such as CDs (Compact Discs) or from servers on the Internet.
[0023] Figure 3 is a block diagram illustrating the hardware configuration of the information processing device 10. As shown in Figure 3, the information processing device 10 includes a processor 1A, memory 2A, input / output interface 3A, peripheral circuitry 4A, and bus 5A. The peripheral circuitry 4A includes various modules. The information processing device 10 does not necessarily have peripheral circuitry 4A. The information processing device 10 may also be composed of multiple physically and / or logically separated devices. In this case, each of the multiple devices may have the above hardware configuration.
[0024] Bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to send and receive data to and from each other. The processor 1A is a processing unit such as a CPU (Central Processing Unit) or GPU (Graphics Processing Unit). Memory 2A is a memory such as RAM (Random Access Memory) or ROM (Read Only Memory). The input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. The input / output interface 3A also includes interfaces for connecting to a communication network such as the Internet. Input devices include, for example, a keyboard, mouse, microphone, physical buttons, touch panel, etc. Output devices include, for example, a display, projection device, speaker, printer, mailer, etc. The processor 1A can issue commands to each module and perform calculations based on their calculation results.
[0025] <Functional Configuration> Next, the functional configuration of the information processing device 10 will be described in detail. Figure 1 is an example of a functional block diagram of the information processing device 10. As shown in the figure, the information processing device 10 has an acquisition unit 11, a speech information generation unit 12, a state information generation unit 13, a detection unit 14, and an output unit 15.
[0026] The acquisition unit 11 acquires the target image.
[0027] The "target image" is an image that is monitored by the monitor. The monitor is a person who monitors the target image and detects the occurrence of problems. The monitor may monitor live images or monitor images taken in the past retrospectively. The target image is a moving image, but it may also be a still image. Furthermore, the target image may be a color image (e.g., an RGB image) or a grayscale image (a black and white image). Furthermore, the target image may be an image created by detecting visible light, or an image created by detecting other electromagnetic waves such as infrared rays.
[0028] The acquisition unit 11 acquires target images generated by a predetermined camera. The acquisition unit 11 can acquire multiple target images generated by each of multiple cameras. The multiple cameras may photograph different locations from each other, or they may photograph the same location from different directions. The cameras may be installed in predetermined locations. In addition, the cameras may be mounted on a mobile device. Examples of mobile devices include, but are not limited to, automobiles, motorcycles, trains, buses, airplanes, ships, drones, and robots.
[0029] The acquisition unit 11 can acquire target images generated by the camera in real time. Alternatively, the acquisition unit 11 may acquire images generated by the camera in batch processing. The acquisition unit 11 may be connected to the camera or a storage device that stores target images generated by the camera in a communicative manner. The acquisition unit 11 may then acquire target images by communicating with these devices. In addition, the acquisition unit 11 may acquire target images input to the information processing device 10 by an operator of the information processing device 10.
[0030] "Acquisition" includes at least one of the following: the device retrieving data or information stored in another device or storage medium (active acquisition), and the device inputting data or information output from another device into its own device (passive acquisition). Examples of active acquisition include making a request to another device and receiving a reply, and accessing and reading data from another device or storage medium. Examples of passive acquisition include receiving information that is delivered (or transmitted, push notification, etc.). Furthermore, acquisition may also involve selecting and acquiring data or information from among the received data or information, or selecting and receiving data or information that has been delivered.
[0031] The speech information generation unit 12 generates speech information regarding the speeches of the people included in the target image. The speech information indicates at least one of the following: the content of the speech, the type of language, and the degree of volume of the voice. The speech information may also include other information related to the speech.
[0032] "Statement content" refers to the content of a statement made by a person included in the target image. In other words, statement content indicates what was said by a person included in the target image.
[0033] "Language type" indicates the type of language spoken by the person in the image. Examples of language types include Japanese, English, etc.
[0034] "Voice volume level" indicates the volume of the voices of the people included in the target image. The voice volume level may be indicated by a predefined classification such as "loud," "somewhat loud," "normal," or "quiet." Although the voice volume level is classified into four levels here, it may be classified into any number of levels. In addition, the voice volume level may be indicated by a widely known unit such as decibels.
[0035] The speech information generation unit 12 can generate speech information by any of the following methods. Method 2 and Method 3 can be adopted when there is voice data collected at the location where the target image was taken. (Method 1) Analyze the target image using lip-reading to generate speech information. (Method 2) Analyze the voice data to generate speech information. (Method 3) Integrate the result of analyzing the voice data and the result of analyzing the target image using lip-reading to generate speech information.
[0036] "Method 1" Regarding Method 1, an explanation will be given. In Method 1, first, the speech information generation unit 12 detects the face of a person from the target image. When the target image includes a plurality of persons, the speech information generation unit 12 detects the faces of the plurality of persons. Then, the speech information generation unit 12 analyzes the image of each face using lip-reading to specify the speech content and language type of each person. The configuration for detecting a person's face from the image and the configuration for analyzing the image using lip-reading can be realized by adopting well-known technologies.
[0037] Furthermore, the speech information generation unit 12 may estimate the degree of loudness of the voice in each speech content. The degree of loudness of the voice may be indicated by a pre-defined classification such as, for example, "loud voice", "somewhat loud voice", "normal", "soft voice", etc. Here, the degree of loudness of the voice is classified into four levels, but it may be classified into other numbers of levels.
[0038] The speech information generation unit 12 can estimate the degree of loudness of the voice based on, for example, the way the mouth opens or the expression. In one example, the speech information generation unit 12 can use a classifier generated by machine learning based on learning data that associates the image of the speaker's face at the time of speech with the degree of loudness of the voice (correct label). <{
[0039] In addition, the speech information generation unit 12 may input the image of the speaker's face at the time of speech into a generative AI (artificial intelligence) and have it estimate the degree of loudness of the voice of the speaker. For example, when the degree of loudness of the voice is classified as "loud voice", "somewhat loud voice", "normal", "soft voice", etc., the speech information generation unit 12 can have it estimate which of the classifications the degree of loudness of the voice of the speaker of the input image corresponds to.
[0040] Generative AI is a concept that includes a large language model (LLM). Generative AI can also include an image analysis model. Generative AI interprets the content of prompts input in natural language, generates responses to the input prompts based on pre-trained data, and outputs responses expressed in natural language. Generative AI can learn from a vast amount of diverse data, such as that publicly available on the internet. Furthermore, Generative AI can accept input such as images along with prompts, interpret the content of the input images, and generate responses that reflect the results of the interpretation. Generative AI can also generate and output images. In this embodiment, the information processing device 10 can input prompts entered by an operator into the Generative AI. Also, in this embodiment, the information processing device 10 can read prompts pre-registered by an operator and input them into the Generative AI. Furthermore, in this embodiment, the information processing device 10 can read prompts from a plurality of pre-registered prompts by an operator according to predetermined rules and input them into the Generative AI. Furthermore, the information processing device 10 of this embodiment can generate prompts by filling predetermined information into the blank spaces of a prompt (template) registered in advance by an operator, and input them into the generating AI. The information processing device 10 of this embodiment may obtain the information to be filled into the blank spaces from another device (such as a server), generate it itself, or accept input from an operator. Furthermore, the information processing device 10 of this embodiment can generate prompts according to predetermined rules and input the generated prompts into the generating AI. The information processing device 10 of this embodiment can input reference information to be referenced along with the prompts into the generating AI. The information processing device 10 of this embodiment can input reference information entered by an operator into the generating AI. Furthermore, the information processing device 10 of this embodiment can obtain reference information from another device (such as a server) and input it into the generating AI. Furthermore, the information processing device 10 of this embodiment can generate reference information according to predetermined rules and input the generated reference information into the generating AI.The generative AI includes, for example, ChatGPT, Gemini, Firefly, Claude, etc., but is not limited thereto. The information processing apparatus 10 of this embodiment may include a generative AI. Alternatively, an external device configured to be communicable with the information processing apparatus 10 of this embodiment may include a generative AI.
[0041] Such generative AI is realized, for example, by a neural network. A neural network includes a plurality of artificial neurons and has synapses connecting each neuron. Each synapse has a weight. When such a neural network receives an input, it performs calculations using the weights associated with each synapse and outputs an output corresponding to the input. A model representing the connection relationship between neurons and synapses is stored in a memory, for example, in the form of software. Alternatively, the model may be realized as a dedicated circuit. Similarly, the weights of each synapse are also stored in a memory in the form of software. Alternatively, a circuit representing the weights may be implemented in a dedicated circuit. When a generative AI is configured using a plurality of models, it is not necessarily required that all the models be stored on the same memory. There are various models using such neural networks. A generative AI may be realized by adopting and replacing various models such as Transformer, convolutional neural network (CNN), recurrent neural network (RNN), etc.
[0042] The speech information generation unit 12 issues person identification information to each person detected in the target image. The speech information generation unit 12 then generates speech registration information, which is linked to each person's identification information and contains the speech information of each person, and can store this information in a predetermined storage device. The speech information generation unit 12 can also manage the position of each detected person within the target image using tracking technology that tracks objects moving within the image. The speech information generation unit 12 may also register information indicating the position of each person within the target image in the speech registration information. The predetermined storage device may be located within the information processing device 10, or it may be located within an external device configured to communicate with the information processing device 10 (the same applies hereinafter).
[0043] Furthermore, the speech information generation unit 12 can determine whether the person detected in each of multiple frame images is the same person, using techniques such as tracking people within images and facial recognition techniques that utilize facial features. When a new person is detected, the speech information generation unit 12 can issue new person identification information linked to that new person.
[0044] Furthermore, the speech information generation unit 12 may register information (such as date and time information) indicating the timing of each person's speech in the speech registration information. The speech information generation unit 12 can identify the timing of each person's speech based on the date and time each target image was taken. The means for identifying the date and time each target image was taken include, but are not limited to, the use of metadata for each target image.
[0045] "Method 2" Method 2 will now be explained. In Method 2, the speech information generation unit 12 can identify the content of the speech, the language type, the volume of the voice, etc., recorded in the speech data by analyzing the speech data using speech analysis technology. In this example, the volume of the voice can be expressed in a widely known unit such as decibels.
[0046] In Method 2, the speech information generation unit 12 may generate speech registration information by simply transcribing the audio data without distinguishing between speakers. The speech information generation unit 12 may further register the language type and volume of each speech in this speech registration information. The speech information generation unit 12 may also register information indicating the timing of each speech (such as date and time information) in the speech registration information. The speech information generation unit 12 can identify the timing of each speech based, for example, on the metadata of the audio data.
[0047] As a variation of Method 2, the speech information generation unit 12 may distinguish the speaker of a speech recorded in the audio data based on a voiceprint or the like. In this case, the speech information generation unit 12 can issue person identification information for each speaker. The speech information generation unit 12 can then generate speech registration information that registers the speech information of each person linked to each person identification information. The speech information generation unit 12 may also register information (such as date and time information) indicating the timing of each speech by each person in the speech registration information. The speech information generation unit 12 can, for example, identify the timing of each speech based on the metadata of the audio data.
[0048] "Method 3" Method 3 will now be explained. In Method 3, the speech information generation unit 12 generates speech registration information using the method of Method 1 and the modified method of Method 2. The speech information generation unit 12 then determines the identity between the person identified by the person identification information issued in Method 1 and the person identified by the person identification information issued in Method 2, based on the similarity of the speech content. The speech information generation unit 12 then links the person identification information that has been determined to be for the same person. As a result, the speech registration information generated in Method 1 and the speech registration information generated in Method 2 concerning the same person are integrated. The calculation of the similarity of the speech content can be achieved using a technique for calculating the similarity of strings. If the similarity between a certain speech content recorded in the speech registration information generated in Method 1 and a certain speech content recorded in the speech registration information generated in Method 2 is above a threshold, the speech information generation unit 12 can determine that the speakers of those speeches are the same person.
[0049] By integrating the utterance registration information generated by Method 1 with the utterance registration information generated by Method 2, the utterance registration information can be enriched by complementing each other's information.
[0050] The state information generation unit 13 generates state information relating to the state of a person included in the target image. The state information indicates at least one of the following: facial expression, emotion, vital information, and behavior.
[0051] "Facial expression" refers to the facial expression of a person in the image. There are various ways to classify facial expressions. For example, facial expressions can include anger, joy, acceptance, surprise, fear, sadness, disgust, and anticipation.
[0052] "Emotion" refers to the feelings of the person depicted in the image. There are various ways to classify emotions. For example, emotions can include anger, joy, acceptance, surprise, fear, sadness, disgust, and anticipation.
[0053] "Vital information" refers to the vital signs of a person included in the image. Examples of vital information include, but are not limited to, heart rate, respiratory rate, blood oxygen saturation, blood pressure, and stress level.
[0054] "Action" refers to an action performed by a person included in the target image. Multiple types of actions related to the problem event detected by the detection unit 14 (described later) may be defined in advance. The state information may indicate that the person included in the target image has performed one of those multiple types of actions. Examples of action types include, but are not limited to, "throwing an object," "kicking," "hitting," "shouting," and "apologizing."
[0055] The state information generation unit 13 detects people from the target image. If the target image contains multiple people, the state information generation unit 13 detects multiple people. If the speech information generation unit 12, as described above, detects people from the target image and issues person identification information, the state information generation unit 13 may acquire the result.
[0056] The state information generation unit 13 generates the state information described above by analyzing the image of each person. The state information generation unit 13 then generates state registration information linked to the person identification information of each person and can store it in a predetermined storage device. By storing the state registration information in a predetermined storage device, past state information can be referenced when the same person is detected again. This allows for confirmation that the person's image was previously a person requiring monitoring.
[0057] The configuration for identifying facial expressions, emotions, vital information, etc., through image analysis can be realized by employing widely known technologies. Furthermore, the state information generation unit 13 can estimate the actions being performed by a person in the target image, for example, using a classifier generated by machine learning. Alternatively, the state information generation unit 13 may input the target image into a generating AI and have it estimate the actions being performed by a person in the target image. For example, the state information generation unit 13 may have the generating AI estimate which of the predefined types of actions the person in the target image is performing.
[0058] The detection unit 14 detects the occurrence of a problem event based on at least one of the speech information and the state information. In one example, the detection unit 14 detects a person requiring monitoring from among the people included in the target image based on the speech information and the state information. The detection unit 14 then detects the detection of a person requiring monitoring as the occurrence of a problem event.
[0059] A "person requiring monitoring" is someone who needs to be monitored. For example, a person who is speaking in an unusual manner would be considered a person requiring monitoring. Examples of people speaking in an unusual manner include, but are not limited to, those who are shouting or threatening, making strange noises, screaming, or speaking at a volume above a certain level.
[0060] The conditions for detecting a person as requiring surveillance (hereinafter referred to as "detection conditions") are defined in advance and stored in a predetermined storage device. The detection unit 14 determines whether each person is a person requiring surveillance based on the speech information and status information of each person detected from the target image.
[0061] Figure 4 shows an example of the detection conditions. In the illustrated detection conditions, if the confidence level of each person's statement that falls into a predetermined category is above a threshold, the person who made that statement is detected as a person requiring monitoring. In the illustrated example, categories such as shouting, screaming, and intimidation are defined as the categories to be judged. And corresponding to each, X 1 ~X 3 A threshold is set for each threshold X 1 ~X 3 These values may be different from each other, or they may be the same.
[0062] In this example, the detection unit 14 calculates the confidence level of each person's statement to belong to each category (shouting, screaming, intimidation, etc.) based on each person's statement information and state information. More specifically, the detection unit 14 can calculate the confidence level of each person's statement to belong to each category (shouting, screaming, intimidation, etc.) based on each person's statement indicated by the statement information and each person's state at the time of each statement indicated by the state information. The detection unit 14 then compares the calculated confidence level with the threshold defined by the detection conditions to determine whether each person's statement satisfies the detection conditions. The detection unit 14 then detects the person who made the statement that satisfies the detection conditions as a person requiring monitoring.
[0063] There are various methods for calculating the reliability score. One example is the use of a generating AI. The detection unit 14 instructs the generating AI to calculate the reliability score for each statement made by each person, based on each statement made by each person, as indicated by the statement information, and the state of each person at the time of each statement, as indicated by the state information, so that each statement by each person belongs to each category (shouting, screaming, intimidation, etc.). At this time, the detection unit 14 may also input target images showing the state of each person at the time of each statement into the generating AI.
[0064] Returning to Figure 4, another example of the detection conditions defines that if the volume of a person's voice is above a threshold, that person is detected as a person requiring monitoring. In this example, the detection unit 14 determines whether the volume of each person's voice at the time of each statement, as indicated in the statement information, is above the threshold. The detection unit 14 then detects any person who speaks at a volume above the threshold as a person requiring monitoring.
[0065] Returning to Figure 1, the output unit 15 outputs the target image and the detection results from the detection unit 14. The output unit 15 can output the above via output devices such as display devices (displays, projection devices, etc.), speakers, and warning lamps.
[0066] For example, the output unit 15 can display the target image on the display device. Alternatively, the output unit 15 may display the detection result from the detection unit 14 on the display device. In addition, the output unit 15 may output the detection result from the detection unit 14 as sound via a speaker. Furthermore, if the detection result from the detection unit 14 indicates the occurrence of a problem event, the output unit 15 may output a warning sound via a speaker. In addition, if the detection result from the detection unit 14 indicates the occurrence of a problem event, the output unit 15 may illuminate a warning lamp.
[0067] An output device is installed in a location where a monitor monitoring the target image can check the output information. The monitor detects the occurrence of a problem based on the information output by such an output device.
[0068] Next, an example of the processing flow of the information processing device 10 will be explained using the flowchart in Figure 2. Note that the purpose here is to explain just one example of the processing flow. Details of each process have been described above, so explanations will be omitted here as appropriate.
[0069] In S10, the information processing device 10 acquires the target image. In S11, the information processing device 10 generates speech information regarding the speech of a person included in the target image. In S12, the information processing device 10 generates state information regarding the state of a person included in the target image. In S13, the information processing device 10 detects the occurrence of a problem event based on at least one of the speech information and the state information. In S14, the information processing device 10 outputs the detection result.
[0070] Note that the processing order of S11 and S12 is not limited to this example. For example, S12 may be performed before S11, or S11 and S12 may be performed in parallel.
[0071] <Effects and Effects> The information processing device 10 of the second embodiment can achieve the same effects and effects as the information processing device 10 of the first embodiment.
[0072] Furthermore, the information processing device 10 can detect the occurrence of a problem event based on speech information and state information. Speech information can indicate at least one of the following: content of speech, language type, and degree of volume. State information can indicate at least one of the following: facial expression, emotion, vital information, and behavior. By using such an information processing device 10 to detect the occurrence of a problem event based on at least one of speech information and state information, the occurrence of a problem event can be detected with high accuracy.
[0073] Furthermore, the information processing device 10 can detect a person who is speaking in an unusual manner as a person requiring monitoring, based on at least one of the speech information and the state information. The information processing device 10 can then detect the detection of a person requiring monitoring as the occurrence of a problem event. By detecting the occurrence of a problem event in this way, the information processing device 10 can detect the occurrence of a problem event with high accuracy.
[0074] <<Third Embodiment>> The information processing device 10 of the third embodiment outputs the target image and the detection results on a characteristic screen. This will be described in detail below.
[0075] The output unit 15 can generate and output screen 1 as shown in Figure 5. When a monitor performs monitoring using multiple screens, the output unit 15 can generate and output a single screen by arranging multiple screens 1 side by side. Alternatively, the output unit 15 may generate multiple screens 1 independently and display each screen 1 on each display device.
[0076] The screen 1 shown in Figure 5 has a display area 2 and a display area 3, which are parts of the screen 1.
[0077] The target image is displayed in display area 2. The output unit 15 may also display a live image (target image) in display area 2. Alternatively, the output unit 15 may play back and display a previously recorded target image in display area 2.
[0078] In the display area 3, the speech history is displayed in chronological order. The output unit 15 displays the speech content of the person included in the target image in the display area 3 in chronological order.
[0079] As shown in FIG. 5, the output unit 15 may display the person identification information of the person detected from the target image in association with each person on the target image (see the display area 2). Further, as shown in FIG. 5, the output unit 15 may display the person identification information of the speaker of each speech in association with each speech indicated by the speech history in the display area 3. "ID: XXX", "ID: YYY", etc. shown in the figure indicate the person identification information. Such display enables the user to grasp which person among the persons included in the target image made each speech indicated by the speech history.
[0080] Further, the output unit 15 may superimpose and display information indicating the speech content of each person included in the target image in text on the target image. In the example of FIG. 5, the output unit 15 displays a balloon indicating the speech content in text in association with the speaker of each speech. Such display enables the user to intuitively grasp the speech content of each person included in the target image.
[0081] Further, the output unit 15 can superimpose and display information for identifying the person to be monitored on the target image. In the example of FIG. 5, frames w 1 and w 2 surrounding the face of each person detected from the target image are displayed. The output unit 15 can identify and display the person to be monitored by making the display mode of the frame w 1 surrounding the face of the person to be monitored different from the display mode of the frame w 2 surrounding the face of other persons. For example, the output unit 15 may make the color, thickness, type, etc. of the line of the frame w 1 surrounding the face of the person to be monitored different from the display mode of the frame w 2 surrounding the face of other persons. Alternatively, the output unit 15 may blink the frame w 1 surrounding the face of the person to be monitored. Alternatively, the output unit 15 may expand the face region of the person to be monitored and superimpose it on the target image. Alternatively, the output unit 15 may expand the face region of the person to be monitored and display it in a separate window.
[0082] Another example of identifying and displaying individuals under surveillance is that the output unit 15 may identify individuals under surveillance by differentiating the display of text indicating the content of their statements from the display of text indicating the content of other individuals' statements. For example, the output unit 15 may differentiate the color, boldness, and size of the text indicating the content of the statements of individuals under surveillance from the display of text indicating the content of other individuals' statements. The text indicating the content of the statements of individuals under surveillance includes at least one of the text displayed in the target image in display area 2 and the text displayed in the statement history in display area 3.
[0083] In addition, the output unit 15 may identify and display persons under surveillance by differentiating the display of the speech bubble indicating the content of the statements of persons under surveillance from the display of the speech bubbles indicating the content of statements of other persons. For example, the output unit 15 may differ the color, border thickness, size, etc., of the speech bubble indicating the content of the statements of persons under surveillance from the display of the speech bubbles indicating the content of statements of other persons. The output unit 15 may always display the speech bubble indicating the content of the statements of persons under surveillance from the display of the speech bubbles indicating the content of statements of other persons, or it may do so only when predetermined conditions are met. The predetermined conditions may be, for example, specific statements (statements containing specific keywords, statements that have been detected to be above a predetermined volume, etc.), but other examples may also be used.
[0084] The information used to identify individuals requiring monitoring is not limited to the examples described above. For example, the output unit 15 may display text information such as "Require Monitoring" or a mark indicating that the individual requires monitoring around the person requiring monitoring.
[0085] This type of display allows users to intuitively identify individuals who require monitoring within the target image.
[0086] The output unit 15 can continue displaying the information identifying the person requiring monitoring, as described above, until a predetermined termination time. One example of a termination time is when a predetermined time has elapsed since the start of displaying the information identifying the person requiring monitoring. Another example of a termination time is when the user (monitoring officer, etc.) gives an instruction to terminate the display. In this way, by continuing to display the information identifying the person requiring monitoring for a certain period of time, the monitoring officer can easily continue tracking / monitoring the person requiring monitoring during that time.
[0087] Furthermore, the output unit 15 may change the display method of the information identifying the person under surveillance according to the number of problematic statements made by the person under surveillance. Also, the output unit 15 may change the display method of the text indicating the statements of the person under surveillance in the statements shown in the statement history of the display area 3 according to the number of problematic statements made by the person under surveillance. For example, as the number of problematic statements increases, the output unit 15 may change the display method of the information identifying the person under surveillance. 1 The display patterns of speech bubbles and text can be changed in stages. In this example, the detection unit 14 can detect problematic statements made by a person under surveillance based on the person's statement information and count the number of such statements. Problematic statements are statements that satisfy the detection conditions described in the second embodiment.
[0088] Furthermore, if a problem event has occurred, the output unit 15 may display information indicating this on screen 1. In the example in Figure 5, the output unit 15 displays the text information "Warning" on the target image in display area 2. Note that the information indicating that a problem event has occurred is not limited to this. For example, a warning icon or warning image may be displayed instead of the text information "Warning".
[0089] Furthermore, the output unit 15 may vary the display format of the speech bubble content displayed in the target image in display area 2, and the display format of the speech history in display area 3, for each individual. For example, the display format of the speech bubble may vary for each individual, or the display format of the text may vary for each individual. In this way, speeches by multiple individuals can be intuitively identified.
[0090] Other configurations of the information processing device 10 can be the same as those in the first and second embodiments.
[0091] The information processing device 10 of the third embodiment can achieve the same effects as the information processing device 10 of the first and second embodiments. Furthermore, the information processing device 10 can output the target image and the detection results on a distinctive screen. With such an information processing device 10, the observer can easily grasp the situation shown in the target image.
[0092] <<Fourth Embodiment>> In the example described in the second embodiment, the information processing device 10 detected "detection of a person requiring monitoring" as "occurrence of a problem event." In the fourth embodiment, the information processing device 10 detects "detection of a person requiring monitoring and the fact that the person engaging in dialogue satisfies the conditions" as "occurrence of a problem event." This will be explained in detail below.
[0093] As described in the second embodiment, the detection unit 14 identifies individuals whose spoken information and status information meet the detection conditions as individuals requiring monitoring.
[0094] The detection unit 14 then identifies the person conversing with the person under surveillance based on at least one of the target image and the speech information. More specifically, the detection unit 14 identifies the person conversing with the person under surveillance at the time the person under surveillance made a problematic statement. The concept of a problematic statement is as described in the third embodiment.
[0095] There are various methods for identifying the person engaging in dialogue. For example, the detection unit 14 may analyze the target image and detect the gaze of the person under surveillance. The detection unit 14 may then identify the person in the line of sight of the person under surveillance as the person engaging in dialogue.
[0096] In addition, the detection unit 14 may analyze the target image and detect the direction of the face of the person to be monitored. The detection unit 14 may then identify the person in the direction the face of the person to be monitored is facing as the person to be spoken to.
[0097] Alternatively, the target image may be analyzed to detect the orientation of the person under surveillance. The detection unit 14 may then identify the person in the direction the person under surveillance is facing as the person engaging in dialogue.
[0098] In addition, the detection unit 14 may identify the person who is conversing with the person under surveillance based on the consistency between the statements made by the person under surveillance and the statements made by each of the other individuals, as well as the pace of the conversation. For example, the detection unit 14 may input the statements made by the person under surveillance and the statements made by each of the other individuals into the generating AI and have it estimate the person who is conversing with the person under surveillance.
[0099] After identifying the person being spoken to, the detection unit 14 determines whether the person being spoken to meets the conditions (hereinafter referred to as "speaker conditions"). If the person under surveillance shouts or makes threats, the speaking person's facial expressions, emotions, vital signs, and actions may reflect these. For example, the speaking person's facial expressions and emotions may be surprise, fear, sadness, disgust, etc. Also, the speaking person's heart rate, respiratory rate, blood pressure, and stress level may increase. Also, the speaking person's blood oxygen saturation may decrease. Also, the speaking person may perform actions such as apologizing.
[0100] The conditions for a person to be identified as a dialogue subject are that the dialogue subject's facial expressions, emotions, changes in vital signs, and actions are consistent with those of a person under surveillance who would yell, threaten, or otherwise behave in such a manner. At least one of the following is defined in advance as a dialogue subject condition: the dialogue subject's facial expressions, emotions, changes in vital signs, and actions. The detection unit 14 determines whether the dialogue subject's facial expressions, emotions, changes in vital signs, and actions satisfy the dialogue subject condition based on at least one of the dialogue subject's utterance information and state information.
[0101] The dialogue participants may include at least one of the following conditions. Note that the following examples are merely illustrative and not limited to these examples. - Facial expression is one of surprise, fear, sadness, or disgust. - Emotion is one of surprise, fear, sadness, or disgust. - Heart rate is above the threshold. - Heart rate is trending upward (e.g., "Most recent measurement is higher than past measurement results," "Most recent measurement is more than the threshold higher than past measurement results," etc.). - Respiratory rate is above the threshold. - Respiratory rate is trending upward (e.g., "Most recent measurement is higher than past measurement results," "Most recent measurement is more than the threshold higher than past measurement results," etc.). - Blood pressure is above the threshold. - Blood pressure is trending upward (e.g., "Most recent measurement is higher than past measurement results," "Most recent measurement is more than the threshold higher than past measurement results," etc.). - Stress level is above the threshold. - Stress level is trending upward (e.g., "Most recent measurement is higher than past measurement results," "Most recent measurement is more than the threshold higher than past measurement results," etc.). - Blood oxygen saturation is below the threshold. - Blood oxygen saturation is trending downward (e.g., "Most recent measurement is lower than past measurement results," "Most recent measurement is more than the threshold lower than past measurement results," etc.). - The action is either an apology or crying.
[0102] The detection unit 14 then detects "the occurrence of a problem event" when "a person requiring monitoring is detected and the person engaging in dialogue meets the conditions for being a dialogue person."
[0103] Furthermore, the detection unit 14 may identify that a person in the conversation has performed a predetermined action, such as an apology, based on the spoken information of the person in the conversation. When performing a predetermined action such as an apology, a predetermined statement such as "I'm sorry" may be made. The detection unit 14 may identify that a person in the conversation has performed a predetermined action such as an apology by detecting characteristic statements made during such predetermined actions.
[0104] Next, an example of the processing flow of the information processing device 10 will be explained using the flowchart in Figure 6. Note that the purpose here is to explain just one example of the processing flow. Details of each process have been described above, so explanations will be omitted here as appropriate.
[0105] In S20, the information processing device 10 acquires the target image. In S21, the information processing device 10 generates statement information regarding the statements made by the person included in the target image. In S22, the information processing device 10 generates state information regarding the state of the person included in the target image. In S23, the information processing device 10 identifies the person who made the problematic statement based on the statement information and state information of the person included in the target image. In S24, the information processing device 10 identifies the person who is interacting with
[0106] In S25, the information processing device 10 utilizes the results of identifying the person to be monitored based on the spoken information and status information of the person to be monitored. The information processing device 10 also determines whether the person to be spoken to satisfies the above-mentioned conditions for the person to be spoken to, based on at least one of the spoken information and status information of the person to be spoken to. If the information processing device 10 detects the person to be monitored and the person to be spoken to satisfies the conditions for the person to be spoken to, it determines that a problem event has occurred.
[0107] Note that the processing order of S21 and S22 is not limited to this example. For example, S22 may be performed before S21, or S21 and S22 may be performed in parallel.
[0108] Other configurations of the information processing device 10 can be the same as those of the first to third embodiments.
[0109] The information processing device 10 of the fourth embodiment can achieve the same effects as the information processing device 10 of the first to third embodiments. Furthermore, the information processing device 10 can detect the occurrence of a problem event based not only on problematic statements made by the person under surveillance, but also on the state of the person interacting with the person under surveillance. By further considering the state of the person interacting with the person under surveillance, the occurrence of a problem event can be detected with greater accuracy.
[0110] <<Modifications>> Below, modifications applicable to the first to fourth embodiments are described. In these modifications, the same effects and advantages as those of the first to fourth embodiments are achieved.
[0111] As shown in Figure 5, the information processing device 10 is configured to display on screen 1 information indicating the content of each person's statements in the target image in text form. In Figure 5, in display area 2, the content of each person's statements in the target image is shown in text form in speech bubbles. Also in Figure 5, in display area 3, the content of each person's statements in the target image is shown in text form as a statement history.
[0112] In a modified example, the information processing device 10 can switch the display / hide of information indicating the content of the statement in text. For example, the information processing device 10 may switch the display / hide of such information in response to user input.
[0113] The information processing device 10 may switch the display / hide of information for each person included in the target image. That is, the output unit 15 may select at least some of the people included in the target image and switch the display / hide of text indicating what the selected people said.
[0114] For example, the information processing device 10 displays frames associated with each person in the display area 2 of Figure 5. 1 and w 2 The information processing device 10 may accept the user's selection of each person by selecting an option or user identification information associated with each person. In addition, the information processing device 10 may accept the user's selection of each person by selecting user identification information for each person or selecting each person's statements in the statement history displayed in display area 3 of Figure 5. The information processing device 10 may also accept input specifying whether or not to display information showing the content of the statements of the specified person in text format.
[0115] Furthermore, the information processing device 10 may display only the statements of the monitored person in text on screen 1, and may not display the statements of other people. In addition, the information processing device 10 may display only the statements of the monitored person and the person engaging in dialogue in text on screen 1, and may not display the statements of other people.
[0116] In one example, the information processing device 10 does not display text indicating the content of each person's statements until a person to be monitored has been detected. Then, when a person to be monitored is detected, the information processing device 10 starts displaying text indicating the content of that person's statements. The information processing device 10 also starts displaying text indicating the content of the statements of the people in the conversation.
[0117] Such an information processing device 10 can suppress privacy issues caused by unnecessarily informing the monitor of the content of what a person in the target image is saying.
[0118] <<Use Case>> The information processing device 10 can be used in situations where customers and store clerks interact, such as in government offices, banks, convenience stores, and pachinko prize exchange counters. In other words, target images taken in such situations can be input to the information processing device 10 and processed. By using it in such situations, it is possible to detect and address problems between customers and store clerks at an early stage.
[0119] It should be noted that the use of the information processing device 10 is not limited to this example. For example, the information processing device 10 can be used in various public places such as parks and roads. That is, target images taken in such places can be input to the information processing device 10 and processed. By using it in such places, it becomes possible to quickly detect and respond to, for example, a person who is making strange noises alone.
[0120] Although this disclosure has been described above with reference to embodiments, this disclosure is not limited to the embodiments described above. Various modifications to the structure and details of this disclosure are possible, which can be understood by those skilled in the art within the scope of this disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0121] Furthermore, the flowchart used in the above explanation shows multiple steps (processes) in sequence. However, the execution order of the steps performed in each embodiment is not limited to the order in which they are described. In each embodiment, the order of the illustrated steps can be changed to the extent that it does not impede the content.
[0122] Some or all of the above embodiments may also be described as follows, but are not limited to the following: 1. An information processing device having: an acquisition means for acquiring a target image; an acquisition means for generating 3. The information processing device according to 3, wherein the output means selects at least some of the people included in the target image and hides the text indicating the content of the selected people's statements in the target image. 5. The information processing device according to any one of 2 to 4, wherein the detection means detects problematic statements by the person under surveillance based on the statement information, and the output means changes the display mode of the information identifying the person under surveillance according to the number of times the person under surveillance has made problematic statements. 6. The information processing device according to any one of 1 to 5, wherein the detection means identifies a person whose statement information and state information satisfy the detection conditions as a person under surveillance, identifies a dialogue person who is conversing with the person under surveillance based on at least one of the target image and the statement information, and detects the occurrence of the problem event based on at least one of the statement information and state information of the person under surveillance and the statement information and state information of the dialogue person. 7. The information processing device according to any one of 1 to 6, wherein the statement information indicates at least one of the content of the statement, the type of language, and the degree of volume of the voice.8. An information processing device according to any one of 1 to 7, wherein the state information is at least one of facial expressions, emotions, vital information, and behavior. 9. An information processing method comprising: one or more computers acquiring a target image; generating statement information relating to statements made by a person included in the target image; generating state information relating to the state of a person included in the target image; detecting the occurrence of a problem event based on at least one of the statement information and the state information; and outputting the target image and the result of the detection. 10. A program that causes a computer to function as: an acquisition means for acquiring a target image; a statement information generation means for generating statement information relating to statements made by a person included in the target image; a state information generation means for generating state information relating to the state of a person included in the target image; a detection means for detecting the occurrence of a problem event based on at least one of the statement information and the state information; and an output means for outputting the target image and the result of the detection.
[0123] Some or all of the appendices 2 to 8, which are dependent on the information processing device described in appendice 1 above, may also be dependent on the information processing method in appendice 9 and the program in appendice 10 in the same dependent relationship as between appendice 1 and appendices 2 to 8. Furthermore, within the scope of not departing from each of the embodiments described above, some or all of the configurations described as appendices can be realized in various hardware, software, various recording means for recording software, or systems.
[0124] This application claims priority based on Japanese Patent Application No. 2025-017375, filed on 5 February 2025, and incorporates all of its disclosures herein.
[0125] 1. Screen 2. Display area 3. Display area 10. Information processing device 11. Acquisition unit 12. Speech information generation unit 13. Status information generation unit 14. Detection unit 15. Output unit 1A. Processor 2A. Memory 3A. Input / Output I / F 4A. Peripheral circuit 5A. Bus
Claims
1. An information processing device comprising: an acquisition means for acquiring a target image; a statement information generation means for generating statement information relating to statements made by a person included in the target image; a state information generation means for generating state information relating to the state of a person included in the target image; a detection means for detecting the occurrence of a problem event based on at least one of the statement information and the state information; and an output means for outputting the target image and the results of the detection.
2. The information processing apparatus according to claim 1, wherein the detection means identifies a person who is a person who satisfies the detection conditions and whose status information is related to the problem event, and the output means superimposes information identifying the person who is a 3. The information processing apparatus according to claim 2, wherein the output means displays information on the target image that indicates the content of each person's statement in text, and identifies and displays text indicating the content of the statement of the person to be monitored and text indicating the content of the statements of the other persons.
4. The information processing apparatus according to claim 3, wherein the output means selects at least some of the people included in the target image and hides the text indicating the content of what the selected people said in the target image.
5. The information processing apparatus according to claim 2, wherein the detection means detects problematic statements made by the person under surveillance based on the statement information, and the output means changes the display mode of the information identifying the person under surveillance according to the number of times the person under surveillance has made problematic statements.
6. The information processing apparatus according to any one of claims 1 to 5, wherein the detection means identifies a person whose statement information and state information satisfy the detection conditions as a person requiring monitoring, identifies a dialogue person who is interacting with the person requiring monitoring based on at least one of the target image and the statement information, and detects the occurrence of the problem event based on at least one of the statement information and state information of the person requiring monitoring and the statement information and state information of the dialogue person.
7. The information processing device according to any one of claims 1 to 6, wherein the utterance information indicates at least one of the content of the utterance, the type of language, and the degree of loudness of the voice.
8. The information processing device according to any one of claims 1 to 7, wherein the state information indicates at least one of facial expressions, emotions, vital information, and behavior.
9. An information processing method comprising one or more computers: acquiring a target image; generating speech information relating to speeches made by a person included in the target image; generating state information relating to the state of a person included in the target image; detecting the occurrence of a problem event based on at least one of the speech information and the state information; and outputting the target image and the detection result.
10. The information processing method according to claim 9, wherein one or more computers identify a person who is a person to be monitored, whose statement information and status information satisfy the detection conditions and who is related to the problem event, and superimposes information identifying the person to be monitored onto the target image.
11. The information processing method according to claim 10, wherein one or more computers display information on the target image that indicates the content of each person's statements in text, and identify and display text indicating the content of statements of the person to be monitored and text indicating the content of statements of other people.
12. The information processing method according to claim 11, wherein one or more computers select at least some of the people included in the target image, and hide the text indicating the content of what the selected people said in the target image.
13. The information processing method according to claim 10, wherein one or more computers detect problematic statements made by the person under surveillance based on the statement information, and change the display manner of the information identifying the person under surveillance according to the number of times the person under surveillance has made the problematic statements.
14. The information processing method according to any one of claims 9 to 13, wherein one or more computers identify a person who satisfies the detection conditions in the statement information and the state information as a person requiring monitoring; identify a person who is interacting with the person requiring monitoring based on at least one of the target image and the statement information; and detect the occurrence of the problem event based on at least one of the statement information and the state information of the person requiring monitoring and the statement information and the state information of the person interacting.
15. A recording medium that stores a program causing a computer to function as: an acquisition means for acquiring a target image; a speech information generation means for generating speech information relating to speeches made by a person included in the target image; a state information generation means for generating state information relating to the state of a person included in the target image; a detection means for detecting the occurrence of a problem event based on at least one of the speech information and the state information; and an output means for outputting the target image and the results of the detection.
16. The recording medium according to claim 15, wherein the detection means identifies a person who is a person who satisfies the detection conditions and whose status information is related to the problem event, and the output means superimposes information identifying the person who is a 17. The recording medium according to claim 16, wherein the output means displays information on the target image that indicates the content of each person's statement in text, and identifies and displays text indicating the content of the statement of the person to be monitored and text indicating the content of the statements of the other persons.
18. The recording medium according to claim 17, wherein the output means selects at least some of the people included in the target image and hides the text indicating the content of what the selected people said in the target image.
19. The recording medium according to claim 16, wherein the detection means detects problematic statements made by the person under surveillance based on the statement information, and the output means changes the display mode of the information identifying the person under surveillance according to the number of times the person under surveillance has made problematic statements.
20. The recording medium according to any one of claims 15 to 19, wherein the detection means identifies a person whose statement information and state information satisfy the detection conditions as a person requiring monitoring, identifies a person who is conversing with the person requiring monitoring based on at least one of the target image and the statement information, and detects the occurrence of the problem event based on at least one of the statement information and state information of the person requiring monitoring and the statement information and state information of the person conversing.