Video search device, video search method, and computer program

The video search device uses emotion estimation to identify and display key scenes in interrogation videos, addressing the challenge of lengthy footage by pinpointing emotional or behavioral changes, thereby enhancing investigation efficiency.

JP7732527B2Active Publication Date: 2025-09-02NEC CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023579905
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-09
Publication Date
2025-09-02
Estimated Expiration
2042-02-09

AI Technical Summary

Technical Problem

Interrogation videos are lengthy, making it time-consuming for prosecutors and police officers to find specific scenes of interest related to suspect behavior or emotions.

Method used

A video search device utilizing emotion estimation technology to detect changes in a person's emotion in a video, associating change-point information with the video, and controlling the display to cue and display target video portions of interest.

Benefits of technology

Facilitates the efficient retrieval of scenes in interrogation videos where suspect behavior or emotions change, reducing the time required for investigation and providing supportive information for users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007732527000001
    Figure 0007732527000001
  • Figure 0007732527000002
    Figure 0007732527000002
  • Figure 0007732527000003
    Figure 0007732527000003
Patent Text Reader

Abstract

In order to retrieve and present a scene that is effective in examining a person using video in which the person appears, this video retrieval device comprises a video analysis unit, a setting unit, and a display control unit. The video analysis unit uses an emotion estimation technique to detect, from a to-be-processed video in which a person is captured, a time point when the feeling of the person has changed, and associates changing-point information indicating the time point with the to-be-processed video. The setting unit uses the changing-point information to determine, in the to-be-processed video, a cueing position for cueing a video portion of interest including the time point when the feeling of the person has changed, and associates cueing-position information indicative of the cueing position with the to-be-processed video. The display control unit uses the cueing-position information to control the displaying of a display device so as to cue and display the video portion of interest in the video.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technique for searching for scenes involving a person from a video in which the person is filmed. [Background technology]

[0002] Interrogations conducted by investigative agencies involve interviewing people suspected of committing a crime (suspects) and those involved in the case (witnesses), and suspect interrogations are conducted behind closed doors. For this reason, in order to prevent unfair interrogations conducted behind closed doors, it is mandatory to audio and video record the entire process of suspect interrogations in cases subject to lay judge trials and cases investigated independently by the prosecutor.

[0003] Patent Document 1 (JP 2017-207877 A) relates to a behavioral analysis technology, and discloses a technology that analyzes student behavior from footage of students filmed in a school and detects the occurrence and signs of problems such as bullying. Patent Document 2 (JP 2020-5014 A) discloses a technology that supports police officers in writing reports about incidents. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2017-207877 [Patent Document 2] Japanese Patent Publication No. 2020-5014 Summary of the Invention [Problem to be solved by the invention]

[0005] Because interrogations often last for long periods of time, the video footage of interrogations (hereinafter referred to as interrogation footage) is also long, and it takes time for prosecutors and police officers to find the parts of the footage (scenes) they want to review.

[0006] The present invention has been devised to solve the above-mentioned problems. That is, a main object of the present invention is to provide a technology for searching for scenes that are effective for investigating people who appear in a video when using the video. [Means for solving the problem]

[0007] In order to achieve the above object, one aspect of the video search device according to the present invention comprises: a video analysis unit that uses emotion estimation technology to detect a time point at which an emotion of a person changes in a video of the person, and associates change-point information indicating the time point with the video; a setting unit that determines a cue position for cueing a target video portion including the time point in the video using the change point information, and associates cue position information representing the cue position with the video; a display control unit that controls the display of the display device using the cue position information to cue the target video portion in the video and display it on the display device; Equipped with.

[0008] In addition, one aspect of the video search method according to the present invention is to By computer, Detecting a time point at which a person's emotion changes from a video of the person using emotion estimation technology, and associating change point information indicating the time point with the video; determining a cue position for cueing a video portion of interest that includes the time point in the video using the change point information, and associating cue position information representing the cue position with the video; The display of the display device is controlled using the cue position information so that the target video portion in the video is cue-located and displayed on the display device.

[0009] Furthermore, in one aspect, the program storage medium according to the present invention comprises: a process of detecting a time point at which an emotion of a person changes from a video of the person using emotion estimation technology, and associating change-point information indicating the time point with the video; a process of determining a cue position for cueing a target video portion including the time point in the video using the change point information, and associating cue position information representing the cue position with the video; a process of controlling the display of the display device using the cue position information to cue the target video portion in the video and display it on the display device; The computer program is stored in the memory. [Effects of the Invention]

[0010] According to the present invention, when using a video to check a person appearing in the video, a technique for searching for a scene that is effective for the check can be provided. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a block diagram illustrating an example of the configuration of a video search device according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a diagram illustrating an example of installation of a photographing device. [Figure 3] FIG. 10 is a diagram illustrating an example of a method for detecting a person from a video. [Figure 4] FIG. 10 is a diagram illustrating an example of a video on which information expressing emotions is superimposed. [Figure 5] FIG. 10 is a diagram illustrating a thumbnail display. [Figure 6] FIG. 10 is a diagram illustrating an example of a video on which information indicating a time point at which an emotion changes is superimposed. [Figure 7] 4 is a flowchart illustrating an example of the operation of the video search device of the first embodiment. [Figure 8] FIG. 10 is a block diagram illustrating an example of the configuration of a video search device according to a second embodiment of the present invention. [Figure 9] FIG. 10 is a diagram illustrating the configuration of a video search device according to a third embodiment of the present invention. [Figure 10] FIG. 10 is a diagram illustrating transcription information. [Figure 11] FIG. 10 is a diagram illustrating an example of a display for accepting a search keyword. [Figure 12] FIG. 10 is a diagram illustrating an example of a display showing search results. [Figure 13] FIG. 1 is a diagram illustrating an example of the minimum configuration of a video search device according to the present invention. [Figure 14] 10 is a flowchart illustrating an example of an operation of the video search device. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0013] First Embodiment 1 is a block diagram illustrating the configuration of a video search device according to a first embodiment of the present invention. In the first embodiment, the video search device 2 is a computer device incorporated into a video viewing system 1 together with a display device 5, and it is assumed that the main users of the video search device 2 will be investigative agencies (such as prosecutors and police officers). The video search device 2 has a function that can search for and present video portions that are likely to be of interest to the user from interrogation videos recorded by investigative agencies.

[0014] That is, as shown in FIG. 1, the video retrieval device 2 of the first embodiment is connected to an input device 4, a display device 5, a speaker 6, and a database 8. The input device 4 is a device that is operated by a user to input information to the video retrieval device 2. Specific examples of the input device 4 include a keyboard and a mouse. The display device 5 is a device that displays images such as videos, text, and photographs on a screen. The speaker 6 is a device that outputs sound. The input device 4 and the display device 5 may be integrated into one device to form a touch panel. The input device 4, the display device 5, and the speaker 6 may also be integrated into one device.

[0015] The database 8 is a storage device equipped with a storage medium for storing data. In the first embodiment, the database 8 stores (preserves) data on interrogation videos. The interrogation video, in this case, refers to video recorded of an interrogation conducted by an investigative agency, and is video recorded of the entire interrogation process from when the suspect enters the interrogation room to when the suspect leaves the interrogation room after being interrogated. FIG. 2 is a diagram showing an example of the installation of a camera (video camera) 40 that records interrogation videos. In the example of FIG. 2, the camera 40 is installed on the ceiling 43 of an interrogation room 42. The camera range of the camera 40 installed in this manner is large enough to capture at least the range of movement of the suspect 44 from the entrance / exit of the interrogation room 42 to where the suspect 44 is seated. The camera 40 is also equipped with a microphone, and has the function of recording audio during the interrogation in addition to video.

[0016] Interrogation video data (video data) captured by such camera device 40 and stored in database 8 is associated with video data identification information that identifies the video data and shooting time information that indicates the date and time of shooting. The video data may also be associated with information that makes it easier to search for the video data from database 8, such as the name of the interrogator who conducted the interrogation or the name of the case. Database 8 also stores data of the audio of the interrogation recorded by camera device 40 (audio data). This data of the audio of the interrogation is associated with data of the interrogation video captured when the audio was recorded. This audio data is associated with audio data identification information that identifies the audio data and recording time information that indicates the date and time of recording.

[0017] As shown in FIG. 1, the video retrieval device 2 includes a control device 20 and a storage device 30. The storage device 30 includes a storage medium for storing data and a computer program (hereinafter also referred to as a program) 31. There are multiple types of storage devices, such as magnetic disk drives and semiconductor memory devices, and there are many types of semiconductor memory devices, such as multiple types of RAM (Random Access Memory) and ROM (Read Only Memory). The type of storage device 30 included in the video retrieval device 2 is not limited to one. Computer devices are often equipped with multiple types of storage devices. Here, the type and number of storage devices 30 included in the video retrieval device 2 are not limited, and a description thereof will be omitted. Furthermore, when the video retrieval device 2 is equipped with multiple types of storage devices 30, they will be collectively referred to as storage devices 30.

[0018] The control device 20 is configured with processors such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). The control device 20 can have various functions based on a program 31 by reading and executing the program 31 stored in the storage device 30. Here, the control device 20 has a video analysis unit 21, a setting unit 22, a display control unit 23, and an audio control unit 24 as functional units for retrieving and presenting video portions (scenes) that are likely to be of interest to the user from the interrogation video. That is, video portions (scenes) that users (prosecutors or police officers) want to view for investigation purposes are likely to be those where the suspect's behavior or emotions change. For this reason, the video search device 2 of the first embodiment has a function of searching the interrogation video for video portions (video portions of interest) that include times when the behavior or emotions of a person in the interrogation video change, and presenting the retrieved video portions (video portions of interest) as video portions of interest.

[0019] That is, the video analysis unit 21 performs video analysis of people appearing in the video to be processed using behavioral analysis technology and emotion estimation technology. The video to be processed here refers to interrogation videos stored in the database 8 that have not yet been analyzed by the video analysis unit 21. The timing at which the video analysis unit 21 analyzes the video to be processed is not particularly limited, but specific examples include the following timings. For example, the timing at which the video analysis unit 21 detects that new interrogation videos have been saved in the database 8 by using information on the saving date and time when the interrogation videos were saved in the database 8 can be cited as an example of the video analysis timing. Another example of the video analysis timing is the timing at which a request to start video analysis of the video to be processed is input to the video retrieval device 2 by the user operating the input device 4.

[0020] Here, we will describe an example of video analysis performed by the video analysis unit 21. First, the video analysis unit 21 detects people from the video to be processed. There are many technologies for detecting people from video, and the person detection technology used here is not limited, so its description will be omitted.

[0021] The video analysis unit 21 also uses emotion estimation technology to estimate the emotions of people detected in the video to be processed based on their facial expressions and movements. Here, emotions refer to the feelings people have toward things, and include various emotions such as joy, sadness, anger, resignation, surprise, disgust, fear, relief, calmness, confusion, tension, familiarity, shame, contempt, admiration, murderous intent, satisfaction, jealousy, regret, desire, anxiety, excitement, impatience, and worry. There are multiple emotion estimation technologies for estimating the emotions of people in video. The emotion estimation technology employed by the video analysis unit 21 is not limited, but examples include emotion estimation technology using AI (artificial intelligence). When emotion estimation technology using AI is employed, an emotion estimation model is pre-stored in the storage device 30 of the video retrieval device 2. The emotion estimation model is a model that receives the video to be processed as input and outputs information representing the emotions of people in the video, and is generated by machine learning. In the first embodiment, the number of types of emotions estimated by the video analysis unit 21 and their contents are set appropriately depending on the performance of the video search device 2, the user's requests, and the like.

[0022] The video analysis unit 21 further associates emotion information representing the estimated emotion with the video to be processed. For example, FIG. 3 is an image diagram schematically illustrating a portion of one frame image constituting the video to be processed. For example, if a person 51 is detected in frame image 50, a person detection area Z representing the image area in which the detected person 51 appears is set in frame image 50. In addition, person detection information is associated with frame image 50. The person detection information is information related to the person detected in the frame image, and here includes information representing the position and size of person detection area Z in frame image 50, as well as person identification information identifying the detected person 51. If multiple people are detected in one frame image, person detection information is generated for each detected person, and the multiple person detection information is associated with one frame image. Furthermore, if the same person is detected to appear in multiple frame images by the tracking process, the person detection information related to the same person associated with those frame images includes common person identification information.

[0023] In the first embodiment, emotion information indicating an emotion estimated by the video analysis unit 21 is included in human detection information related to a person estimated to have that emotion, and is associated with the video (frame image) to be processed. That is, in the first embodiment, the human detection information associated with the frame image includes information indicating the position and size of the human detection area Z, human identification information identifying the detected person 51, and information indicating the estimated emotion of the detected person 51.

[0024] The video analysis unit 21 further detects the time point at which the detected human emotion changes from the video to be processed. Furthermore, the video analysis unit 21 associates change-point information indicating the detected time point with the video to be processed. The change-point information may be any information that can identify the time point at which the detected human emotion changed in the video to be processed, and may be represented, for example, by a frame number that identifies the frame image at which the change in emotion was detected. Alternatively, if the video to be processed includes information on the time of shooting, the change-point information may be information on the time corresponding to the time at which the human emotion changed. The change-point information may, for example, be in the form of a list.

[0025] Furthermore, the video analysis unit 21 detects the time point at which the behavior of the person detected from the video to be processed changes using a behavior analysis technique. There are many behavior analysis techniques for detecting human behavior from video, and the behavior analysis technique used here is not limited, so its description will be omitted. Furthermore, the video analysis unit 21 associates behavior information indicating the time point at which the detected behavior changed with the video to be processed. In other words, the behavior information is included in human detection information related to the person corresponding to the information, and is associated with the video (frame image) to be processed.

[0026] Note that the processes of analyzing the behavior of a person detected from the video to be processed and detecting the time point when the behavior changes may be omitted for reasons such as reducing the processing load. Furthermore, in the case of interrogation video, since the suspect's seated position is fixed, it is easy to identify the suspect from the interrogation video even if multiple people are captured in the video. Since the suspect is often the person of interest in interrogation video, the video analysis unit 21 may perform at least the emotion estimation process, which estimates the emotion of only the suspect detected from the interrogation video, and the behavior analysis process. Settings such as whether or not to perform the behavior analysis described above and whether or not to limit the analysis to the suspect are appropriately set according to the processing capabilities of the video search device 2 and the user's requests.

[0027] The setting unit 22 uses the change point information and behavior information acquired by video analysis to determine a cue position for cueing a scene of interest in the video to be processed analyzed by the video analysis unit 21. As described above, a scene of interest is a video portion including a point in time when the behavior or emotion of a person shown in the video changes. Here, the time range of the scene of interest in the video may be appropriately set, taking into consideration that the scene is one that the user is likely to want to watch, as long as it includes a point in time when the person's behavior or emotion changed. For example, an example of the start position (i.e., cue position) of the scene of interest may be a point that goes back a predetermined time (e.g., 10 seconds, or the number of frames equivalent to the set time) from the point in time when the behavior or emotion changed. Note that the video search device 2 may have a function that allows the user to variably set the time to go back. Furthermore, the end position of the scene of interest is determined by the user, and such information does not need to be set in the interrogation video.

[0028] The setting unit 22 further associates cue position information indicating the determined cue position with the video to be processed.

[0029] In the first embodiment, the video processed by the video analysis unit 21 and the setting unit 22 is stored in the database 8 in a state in which the person detection information, change point information, behavior information, and cue position information are associated with each other.

[0030] The display control unit 23 controls the display of the display device 5. One example of display control by the display control unit 23 is the following display control. For example, when a user uses the input device 4 to specify an interrogation video to be viewed, the video retrieval device 2 reads the interrogation video identified by the input information from the database 8. For example, when the display control unit 23 detects that a command to play the interrogation video read from the database 8 has been input to the video retrieval device 2 by the user operating the input device 4, the display control unit 23 plays the interrogation video. The method for playing the interrogation video from the beginning or the method for fast-forwarding or rewinding the interrogation video are not limited here, and therefore, a description thereof will be omitted. Furthermore, when the interrogation video is read from the database 8, audio data associated with the interrogation video is also read from the database 8. The audio control unit 24 controls the audio output from the speaker 6 by synchronizing the audio of the read audio data with the video played by the display control unit 23. The method for controlling this audio output is not limited here, and therefore, a description thereof will be omitted. In addition, if the video search device 2 is equipped with a mute function and the mute mode is set by the user operating the input device 4, the audio control unit 24 will not output audio while the interrogation video is being played back.

[0031] As another display control technique, the display control unit 23 may superimpose analysis information on the interrogation video being played back. The analysis information is information obtained through video analysis by the video analysis unit 21, and in the first embodiment, includes emotional information contained in the human detection information and attitude information representing behavior. FIG. 4 illustrates a specific example in which analysis information is superimposed on the interrogation video being played back. In the example of FIG. 4, the analysis information, which is emotional information (e.g., nervousness) and attitude information (e.g., fidgeting), for the suspect 44 is superimposed on the interrogation video using text. Note that, in the example of FIG. 4, the emotional information and attitude information are expressed using text; however, one or both of the emotional information and attitude information may be expressed using marks or symbols instead of text. Furthermore, text representing the emotional information or attitude information may be written alongside the marks or symbols representing the emotional information or attitude information. Furthermore, in the example of FIG. 4, analysis information for the suspect 44 is superimposed on the interrogation video. Furthermore, if emotional information or attitude information of a person other than the suspect (for example, a police officer) appearing in the interrogation video is associated with the interrogation video through video analysis, the emotional information or attitude information of the person other than the suspect may also be superimposed on the interrogation video. Note that even when the interrogation video is played back in the information display mode in this way, the audio output is controlled by the audio control unit 24 in the same manner as described above.

[0032] Here, a display mode in which emotional information or attitude information is superimposed on the interrogation video being played back is referred to as an information-present mode, and a display mode in which emotional information or attitude information is not superimposed on the interrogation video being played back is referred to as an information-absent mode. The video retrieval device 2 may have a function that allows a user to alternatively select either an information-present mode or an information-absent mode for playing back the interrogation video and accept the selected information. The display control unit 23 may superimpose or not superimpose emotional information or attitude information on the interrogation video being played back according to the selected display mode. Furthermore, the video retrieval device 2 may be configured to allow the display control unit 23 to select either or both emotional information and attitude information as information to be superimposed on the interrogation video being played back. In other words, it is assumed that there are cases in which it is better to display emotional information or attitude information and cases in which it is better not to display it. For this reason, by being able to select whether or not to display information as described above, information can be displayed according to the user's needs, thereby improving the convenience of the video retrieval device 2.

[0033] Furthermore, in the first embodiment, the display control unit 23 has a function capable of playing back interrogation videos in the following recommended search mode. The recommended search mode here is a mode in which a noteworthy scene video is cue-searched and displayed from the interrogation video using cue position information associated with the interrogation video. For example, when the interrogation video includes multiple noteworthy scene videos, the display control unit 23 may cue-search and display the noteworthy scene video by displaying multiple noteworthy scene videos 53 as thumbnails, as shown in FIG. 5 . The thumbnail display here refers to a display in which multiple noteworthy scene videos 53 are lined up and played back from a cue position. When playing back noteworthy scene videos 53 using this thumbnail display, the display control unit 23 starts playing back each noteworthy scene video 53 from its cue position after detecting that the user has operated the input device 4 to instruct the start of playback for each noteworthy scene video 53. Alternatively, the display control unit 23 may start playing back each noteworthy scene video 53 from its cue position without waiting for a command from the user. Such control of the start of playback of the attention scene video 53 may be set in advance, or may be set by the user operating the input device 4. When multiple attention scene videos 53 are played back simultaneously using thumbnail display, the audio control unit 24 does not output audio, for example.

[0034] Another method by which the display control unit 23 cue-ups and displays a scene of interest is skip playback. Skip playback here refers to a display in which, when the skip playback is requested by the user operating the input device 4, the video is skipped from the video portion of the interrogation video currently being played or about to be played back to the cue position of the nearest scene of interest video. The direction in which to jump to the cue position of the scene of interest video is specified by the user, for example, as either a fast-forward direction or a rewind direction.

[0035] Another method by which the display control unit 23 cue-ups and displays the scene video of interest is to display a list of cue positions for the scene video of interest, and play the scene video of interest at the cue position selected from the list.

[0036] As described above, there are a plurality of methods for cueing and displaying a scene of interest. The display control unit 23 may display the scene of interest using one of the plurality of display methods that is set in advance, or may display the scene of interest using a method selected by the user from the plurality of display methods, and an appropriate display method is set.

[0037] Furthermore, while the interesting scene video is being played back, the display control unit 23 refers to the change-point information and behavior information associated with the currently played interrogation video, and displays information indicating the time point at which the suspect's emotions or behavior changed, superimposed on the interrogation video. Fig. 6 is a diagram showing an example of information to be superimposed on the interrogation video to indicate the time point at which the emotions of the suspect 44 changed. Note that Fig. 6 is just one example, and the manner of the information to be superimposed on the interrogation video to indicate the time point at which the emotions or behavior of the suspect 44 changed may be set appropriately taking into consideration the ease of viewing, etc.

[0038] The video retrieval device 2 of the first embodiment is configured as described above. Next, an example of the operation related to video retrieval in this video retrieval device 2 will be described with reference to Fig. 7. Fig. 7 is a flowchart illustrating an example of the operation related to video retrieval in the video retrieval device 2.

[0039] For example, when the video analysis unit 21 in the video retrieval device 2 detects that it is time to analyze the video, it reads the video to be processed (interrogation video) from the database 8 and performs a predetermined video analysis (step 101 in FIG. 7). As a result, the emotions of people such as suspects appearing in the video to be processed are estimated, and the time points at which their emotions or behaviors change are detected. Furthermore, emotion information indicating the estimated emotions, change-point information indicating the time points at which their emotions change, behavior information indicating the time points at which their behaviors change, and the like are associated with the video to be processed (interrogation video).

[0040] Furthermore, the setting unit 22 determines the cue position of the interesting scene video using information on the time point when the emotion or behavior detected by the video analysis unit 21 changed (step 102). The setting unit 22 associates cue position information indicating the determined cue position with the video to be processed. In this way, the video to be processed (interrogation video) associated with the information generated by the setting unit 22 or the information generated by the video analysis unit 21 is overwritten and stored in the database 8, for example.

[0041] Thereafter, when the user operates the input device 4 to specify the interrogation video that the user wishes to view, the video search device 2 reads out the interrogation video identified by the input information from the database 8. Then, the display control unit 23 controls the display of the display device 5 to display the read out interrogation video on the display device 5 (step 103). For example, when the user sets the display in the recommended search mode, the display control unit 23 references the cue position information associated with the interrogation video to cue the attention scene video (attention video portion) and plays back the attention scene video.

[0042] The video search device 2 of the first embodiment is configured as described above and can cue and display (play) a noteworthy scene video (noteworthy video portion) that includes a point in time when the emotions or behavior of the suspect 44 shown in the interrogation video change. As described above, it is considered that the video portions (scenes) that users (prosecutors or police officers) want to view for investigation purposes from the interrogation video are often the portions where the suspect's behavior or emotions change. For this reason, the video search device 2 can search for and display video portions that are expected to be highly necessary for the user to view. In other words, when using video to investigate a person shown in the video, the video search device 2 can search for and provide scenes that are useful for the investigation. This allows the video search device 2 to shorten the time it takes for a user to investigate (search) the interrogation video.

[0043] In addition, in the first embodiment, the video search device 2 can superimpose information obtained by video analysis onto the video, thereby providing information that can assist the user in investigating a suspect when investigating the suspect from interrogation video.

[0044] Second Embodiment A second embodiment of the present invention will be described below. In the description of the second embodiment, parts with the same names as those of the video viewing system described in the first embodiment will be assigned the same reference numerals, and duplicate descriptions of the common parts will be omitted.

[0045] 8 is a block diagram schematically showing the configuration of a video retrieval device 2 according to the second embodiment. In addition to the configuration of the first embodiment, the video retrieval device 2 according to the second embodiment is provided with an audio analysis unit 25 as a functional unit of the control device 20. Like the other functional units, this audio analysis unit 25 is also realized by a processor executing a computer program.

[0046] The audio analysis unit 25 analyzes audio data of the interrogation associated with the interrogation video and estimates the emotion of the person who uttered the recorded audio (also referred to as the speaker) using emotion estimation technology. The timing of the audio analysis by the audio analysis unit 25 is not particularly limited, but one example is timing that coincides with the analysis of the interrogation video by the video analysis unit 21. There are multiple emotion estimation technologies for estimating the emotion of a speaker from audio. The emotion estimation technology employed by the audio analysis unit 25 is not limited, but may include, for example, emotion estimation technology using AI technology. When emotion estimation technology using AI technology is employed, an emotion estimation model for audio is pre-stored in the storage device 30 of the video retrieval device 2. The emotion estimation model for audio is a model that receives audio data to be processed as input and outputs information representing the emotion of the speaker contained in the audio data, and is generated by machine learning. In the second embodiment, the number and content of the types of emotions estimated by the audio analysis unit 25 are appropriately set depending on the performance of the video retrieval device 2, the user's requests, and the like.

[0047] The voice analysis unit 25 further uses the estimated emotion information to detect the time point at which the speaker's emotion contained in the voice data changes. Furthermore, the voice analysis unit 25 associates voice emotion information indicating the estimated emotion and voice change point information indicating the time point at which the speaker's emotion changed with the voice-analyzed voice data. Note that since the interrogation voice contains the voices of multiple people (speakers), the voices of each speaker are distinguished and analyzed in the voice analysis. As a result, the voice emotion information and voice change point information each contain voice identification information that identifies the voice (speaker). The interrogation voice data associated with such voice emotion information and voice change point information is overwritten and stored in the database 8. Note that since the voice data is played back in synchronization with the interrogation video, the voice emotion information and voice change point information are generated, for example, taking synchronization with the video into consideration.

[0048] In the second embodiment, in an interrogation video associated with audio data that has undergone audio analysis, a video portion including a time point at which the speaker's emotion detected from the audio changes is also considered to be a scene of interest video (video portion of interest). Thus, in the second embodiment, the setting unit 22 uses audio change point information associated with the audio data of the interrogation to determine a cue position of the scene of interest video including a time point at which the emotion detected from the audio changes in the interrogation video associated with the audio data. The method for determining this cue position is not limited, but for example, the cue position of the scene of interest video is determined by audio analysis in the same way as the cue position of the scene of interest video in the interrogation video is determined, taking into consideration synchronizing the audio of the interrogation with the interrogation video. In this way, cue position information including the cue position of the scene of interest video determined by audio analysis and the cue position of the scene of interest video determined by video analysis is associated with the interrogation video by the setting unit 22.

[0049] The display control unit 23 controls the display of the interrogation video in the recommended search mode by referencing the cue position information associated with the interrogation video by the setting unit 22, as in the first embodiment. It is assumed that the cue position of the scene of interest video determined using the results of video analysis of the interrogation video and the cue position of the scene of interest video determined using the results of audio analysis of the interrogation audio may be the same. For example, if the interval between the cue position of the scene of interest video determined using the results of video analysis and the cue position of the scene of interest video determined using the results of audio analysis is short, such as a predetermined time width (e.g., 1.5 seconds) or less, the cue position may be set according to a predetermined rule. For example, one possible rule is that the cue position determined using the results of video analysis or the audio analysis is adopted, whichever is earlier in time. Another possible rule is that the midpoint between the cue position determined using the results of video analysis and the audio analysis is adopted as the cue position. Multiple rules are possible, and an appropriate rule is adopted from among these multiple rules.

[0050] Furthermore, the display control unit 23 may refer to audio emotion information representing emotions estimated by audio analysis, and superimpose the information representing emotions estimated by audio analysis on the interrogation video while the interrogation video is being played back. In this case, the display control unit 23 may superimpose the information representing emotions on the interrogation video using a display format such as color coding so that emotions estimated by audio analysis and emotions estimated by video analysis can be distinguished.

[0051] The configuration of the video retrieval device 2 of the second embodiment other than the above-mentioned configuration is the same as the configuration of the video retrieval device 2 of the first embodiment.

[0052] The video retrieval device 2 of the second embodiment is configured as described above, and has the same configuration as the video retrieval device 2 of the first embodiment, and therefore can achieve the same effects as the first embodiment. Furthermore, the video retrieval device 2 of the second embodiment can also provide emotion information estimated by voice analysis, thereby increasing the amount of information that supports the user's operation when investigating a suspect using interrogation video.

[0053] Third Embodiment A third embodiment of the present invention will be described below. In the description of the third embodiment, parts with the same names as the components of the video viewing system described in the first or second embodiment will be assigned the same reference numerals, and duplicate descriptions of the common parts will be omitted.

[0054] In addition to the configuration of the first or second embodiment, the video retrieval device 2 of the third embodiment includes a transcription unit 26 as shown in Fig. 9 as a functional unit of the control device 20. Like the other functional units, this transcription unit 26 is realized by a processor executing a computer program. Note that Fig. 9 does not illustrate the storage device 30 in the video retrieval device 2, or the video analysis unit 21, setting unit 22, display control unit 23, audio control unit 24, and audio analysis unit 25 in the control device 20.

[0055] The transcription unit 26 uses speech recognition technology to recognize the content of the interrogation audio data associated with the interrogation video to be processed and converts it into text data. Furthermore, the transcription unit 26 detects the time when the audio was uttered and generates transcription information in which utterance time information indicating the time when the audio was uttered is associated with text data indicating the content of the audio. Furthermore, the transcription unit 26 associates the generated transcription information with the interrogation video to be processed. Figure 10 is an image diagram that schematically shows the transcription information. In the example of Figure 10, utterance time information 54 indicating the time when the audio was uttered is represented by the elapsed time from the start of the interrogation video.

[0056] In addition, when voice change point information indicating the time point at which the emotion of the speaker who uttered the voice changed is associated with the interrogation video by the voice analysis unit 25 as described in the second embodiment, the transcription unit 26 may reflect the voice change point information in the transcription information. In the example of Fig. 10, the time point indicated by the voice change point information is associated with time information indicating the time point based on utterance time point information 54 and is indicated by a mark 55.

[0057] Furthermore, since the speaker can be distinguished by the voice analysis performed by the voice analysis unit 25 as described in the second embodiment, voice identification information for distinguishing (identifying) the speaker may be associated with the text data representing the voice content, for example, for each sentence.

[0058] In the third embodiment, the video search device 2 further includes a function of searching the text data generated by the transcription unit 26 for words or sentences input by the user operating the input device 4 as search keywords.

[0059] In the third embodiment, the display control unit 23 controls the display of the transcription information generated by the transcription unit 26 in text form on the display device 5. When the transcription information is displayed on the display device 5, it may be possible to select a mode in which the text information representing the audio content is simply arranged in chronological order, and a mode in which the text information representing the audio content classified by speaker based on the audio identification information is arranged in chronological order for each speaker.

[0060] Furthermore, the display control unit 23 controls the display of the display device 5 to accept keywords to be searched (hereinafter also referred to as search keywords), as shown in FIG. 11 . The display control unit 23 also controls the display of search results based on the search keywords in the text data of the transcription information on the display device 5. The method of displaying search results from text data based on the search keywords is not limited here, and the display control unit 23 controls the display of the search results using a predetermined display method. When displaying the search results, if a word or sentence corresponding to the search keyword is searched for in the text data, a video portion corresponding to the time when the searched word or sentence was uttered may be extracted from the interrogation video and displayed on the display device 5 by the display control unit 23. In this case, the text information representing the audio content and the interrogation video corresponding to the search result are displayed side by side, as shown in FIG. 12 , and the interrogation video and audio are played. Alternatively, information (hyperlink) for reading the interrogation video corresponding to the word or sentence corresponding to the search keyword may be embedded in the word or sentence corresponding to the search keyword, and the interrogation video may be played in a pop-up display.

[0061] The configuration of the video retrieval device 2 of the third embodiment other than the above is the same as the configuration of the video retrieval device 2 of the first or second embodiment.

[0062] The video search device 2 of the third embodiment has a configuration similar to that of the video search device 2 of the first or second embodiment, and can therefore achieve the same effects as those of the video search device 2 of the first or second embodiment.

[0063] Furthermore, the video search device 2 of the third embodiment is configured to be able to transcribe audio and display it on the display device 5, thereby increasing the types of information that can be acquired from interrogation videos. In other words, the video search device 2 of the third embodiment can further increase the amount of information that supports the user in their investigations.

[0064] <Other embodiments> The present invention is not limited to the first to third embodiments, and various embodiments can be adopted. For example, in the first to third embodiments, the configuration of the video retrieval device 2 is described using an interrogation video as an example of the video to be processed, but the video to be processed is not limited to an interrogation video, and the video retrieval device according to the present invention can be applied to any video in which people are filmed. For example, the video retrieval device may process a video of a meeting as the video to be processed.

[0065] Fig. 13 is a block diagram illustrating an example of the minimum configuration of a video retrieval device according to the present invention. The video retrieval device 10 shown in Fig. 13 is, for example, a computer device, and by executing a pre-provided computer program, is provided with a video analysis unit 11, a setting unit 12, and a display control unit 13 as functional units.

[0066] The video analysis unit 11 uses emotion estimation technology to detect a time point at which a person's emotion changes in a video to be processed, and associates change-point information indicating the time point with the video to be processed. The setting unit 12 uses the change-point information to determine a cue position for cueing a video portion of interest that includes the time point at which the person's emotion changes in the video to be processed, and associates cue position information indicating the cue position with the video to be processed.

[0067] The display control unit 13 uses the cue position information to control the display of the display device 15 so that the display device 15 cue-ups and displays the target video portion in the video.

[0068] Next, an example of the operation of the video search device 10 related to video search will be described with reference to Fig. 14. For example, when the video search device 10 detects that it is time to process a predetermined video, the video analysis unit 11 uses emotion estimation technology to detect a point in time when the emotion of a person changes in the video to be processed, in which the person is captured. Then, the video analysis unit 11 associates change-point information indicating the detected point in time with the video to be processed (step 201 in Fig. 14).

[0069] Thereafter, the setting unit 12 determines a cue position for cueing a target video portion including a time point at which a person's emotion changes in the video to be processed, using the change point information (step 202). Furthermore, the setting unit 12 associates cue position information indicating the determined cue position with the video to be processed.

[0070] Thereafter, for example, when a request is made to play a video portion of interest in a video associated with cue position information, the display control unit 13 controls the display of the display device 15 to cue and display the video portion of interest using the cue position information (step 203).

[0071] The video search device 10 shown in Fig. 13 can cue up and display (play) a video segment of interest that includes a point in time when a person's emotions change. This allows the video search device 10 to search for and provide scenes that are useful for searching for people who appear in a video. This reduces the time required for video search compared to when a user searches for a video segment of interest.

[0072] The present invention has been described above using the above-described embodiments as exemplary examples. However, the present invention is not limited to the above-described embodiments. In other words, the present invention can be applied in various aspects that can be understood by a person skilled in the art within the scope of the present invention. [Explanation of symbols]

[0073] 2,10 Video search device 5,15 Display device 11,21 Video Analysis Section 12,22 Setting section 23,13 Display control unit 25 Audio analysis section 26 Transcription Department

Claims

1. a video analysis means for detecting a time point at which an emotion of a person changes from a video of the person using an emotion estimation technique, and associating change-point information indicating the time point with the video; a setting means for determining a cue position for cueing a target video portion including the time point in the video using the change point information, and associating cue position information representing the cue position with the video; a display control means for controlling the display of the display device to cue up the target video portion in the video using the cue position information and display the video on the display device; a voice analysis means for detecting, from the voice recorded at the same time as the video is shot, a time point at which the emotion of the speaker who uttered the voice changes using an emotion estimation technique, and associating voice change point information indicating the time point with the video; Equipped with a video portion including a point in time when the emotion of the speaker changes is also set as the video portion of interest, and the setting means further uses the audio change point information to determine the cue position of the video portion of interest in the video, and associates the cue position information representing the cue position with the video; When an interval between a cue position based on a time point at which the person's emotion changed, detected by the video analysis means, and a cue position based on a time point at which the speaker's emotion changed, detected by the audio analysis means, is equal to or shorter than a preset time width, the display control means sets, according to a preset rule, a single position obtained by integrating the cue position based on the time point at which the person's emotion changed and the cue position based on the time point at which the speaker's emotion changed. Video search device.

2. the video analysis means estimates an emotion of the person captured in the video, and associates emotion information representing the estimated emotion with the video; The display control means superimposes information representing the emotion of the person appearing in the video on the video being played back on the display device, using the emotion information. The video search device according to claim 1 .

3. the voice analysis means estimates the emotion of the speaker and associates voice emotion information representing the estimated emotion with the video; The display control means superimposes information representing the emotion of the speaker on the video being played back on the display device by using the voice emotion information.

3. The video search device according to claim 1.

4. The video recording system further includes a transcription unit that converts the audio recorded when the video is shot into text data using a voice recognition technology, generates utterance time information indicating the time when the audio was uttered, and associates transcription information including the text data and the utterance time information with the video; The display control means further controls a display on the display device that accepts keywords to be searched, and also controls a display on the display device that shows search results based on the keywords in the text data.

4. The video search device according to claim 1.

5. The display control means controls at least one of skip playback for skipping the video to the cue position in the video and playing it back, and, when a plurality of pieces of cue position information are associated with the video, thumbnail display for displaying video portions starting from the plurality of cue positions side by side.

5. A video search device according to claim 1.

6. The display control means controls the display of the video being played in one mode selected from an information-present mode in which the information representing the emotion is superimposed on the video being played, and an information-absent mode in which the information representing the emotion is not superimposed on the video being played.

4. The video search device according to claim 2 or 3.

7. By computer, Detecting a time point at which a person's emotion changes from a video of the person using emotion estimation technology, and associating change point information indicating the time point with the video; determining a cue position for cueing a video portion of interest that includes the time point in the video using the change point information, and associating cue position information representing the cue position with the video; A time point at which the emotion of a speaker who uttered the voice changed is detected from the voice recorded at the time the video was shot using emotion estimation technology, and voice change point information indicating the time point is associated with the video; a video portion including a point in time when the emotion of the speaker changes is also set as the video portion of interest, and the cue position of the video portion of interest in the video is determined using the audio change point information, and the cue position information representing the cue position is associated with the video; using the cue position information to cue the video portion of interest in the video and display it on the display device; When controlling the display of the display device, if an interval between a cue position based on a time point at which the person's emotion changed, detected using the video and the emotion estimation technology, and a cue position based on a time point at which the speaker's emotion changed, detected using the audio and the emotion estimation technology, is equal to or less than a preset time width, a single position obtained by integrating the cue position based on the time point at which the person's emotion changed and the cue position based on the time point at which the speaker's emotion changed is set as a cue position according to a preset rule. How to search for videos.

8. a process of detecting a time point at which an emotion of a person changes from a video of the person using emotion estimation technology, and associating change-point information indicating the time point with the video; a process of determining a cue position for cueing a target video portion including the time point in the video using the change point information, and associating cue position information representing the cue position with the video; a process of detecting, using emotion estimation technology, a point in time when the emotion of a speaker who uttered the voice changed from the voice recorded at the time the video was shot, and associating voice change point information indicating the point in time with the video; a process of determining a cue position of the video portion of interest in the video using the audio change point information, and associating the cue position information representing the cue position with the video; a process of controlling the display device to cue up the target video portion in the video using the cue position information and display the video on the display device; a process of controlling display of the display device, when an interval between a cue position based on a time point at which the person's emotion changed, detected using the video and the emotion estimation technology, and a cue position based on a time point at which the speaker's emotion changed, detected using the audio and the emotion estimation technology, is equal to or shorter than a predetermined time width, setting, as a cue position, a single position obtained by integrating the cue position based on the time point at which the person's emotion changed and the cue position based on the time point at which the speaker's emotion changed, in accordance with a predetermined rule; A computer program that causes a computer to execute the following.

Citation Information

Patent Citations

  • Broadcast recording and reproducing apparatus

    JP2006245907A

  • Database system of call center, and its information management method and information management program

    JP2009175336A

  • Video recorder, reproducer and server device

    JP2012227760A

  • Imaging apparatus

    JP2016122945A

  • Behavioral analysis device and program

    JP2017207877A