Person identification device, meeting minutes creation system, person identification method, and program
Patent Information
- Application Number
- JP2025028842
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2026-09-07
AI Technical Summary
【0006】 本発明によれば、カメラが取得した画像情報から発生音を発した人物を特定する処理の負荷を軽減できる。
Smart Images

Figure 2026142010000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a person identification apparatus, a minutes creation system, a person identification method, and a program.
Background Art
[0002] Conventionally, minutes creation assistance technology for assisting a person in charge in creating meeting minutes is known. For example, Patent Document 1 discloses a minutes creation assistance system including: a sound input unit capable of inputting sound uttered during a meeting; a speech recognition unit that recognizes speech and converts it into character data when the sound input unit inputs, as sound, speech uttered by a meeting attendee; a speaker identification unit that identifies an attendee who uttered the speech as a speaker when the sound input unit inputs the speech; and a recording unit that records the character data converted by the speech recognition unit in association with the speaker identified by the speaker identification unit.
Summary of the Invention
Problem to be Solved by the Invention
[0003] In cases where real-time performance is not required, such as when creating minutes after a meeting, there is no problem even if it takes a long time to process identifying the attendee who uttered the speech as the speaker. However, when real-time performance is required, such as when displaying the results of processing that converts speech into character data through recognition and identifies the attendee who uttered the speech as the speaker during the meeting, it is necessary to perform the speaker identification process in a short time. There is a trade-off relationship between speaker identification accuracy and the time taken until a speaker can be identified. In order to accurately identify a person such as a speaker from video information, it is generally necessary to perform a large amount of analysis processing on enormous amounts of audio information and video information, which poses a problem of long processing time. Patent Document 1 does not disclose any description regarding such a problem.
[0004] An object of an embodiment of the present invention is to provide a person identification apparatus that reduces the processing load for identifying a person who uttered sound from image information acquired by a camera.
Means for Solving the Problem
[0005] A person identification device according to one embodiment of the present invention includes a sound detection unit that detects a sound and the direction from which the sound is generated from audio information acquired by a microphone, and an image analysis unit that analyzes image information acquired by a camera. When the sound detection unit detects the sound and the direction from which the sound is generated, the image analysis unit analyzes image information within a predetermined range from the direction of generation as the analysis range to identify the person emitting the sound. [Effects of the Invention]
[0006] According to the present invention, the processing load required to identify the person who emitted the sound from the image information acquired by the camera can be reduced. [Brief explanation of the drawing]
[0007] [Figure 1] This is a diagram illustrating an example of a meeting minutes creation system according to this embodiment. [Figure 2] This is a hardware configuration diagram of an example of a person identification device according to this embodiment. [Figure 3] This is a hardware configuration diagram of an example computer. [Figure 4] This is a functional configuration diagram of an example of a meeting minutes creation system according to this embodiment. [Figure 5] This is an illustrative diagram showing an example of the processing performed by the person tracking unit. [Figure 6] This is an illustrative diagram showing an example of the processing performed by the participant management department. [Figure 7] This is an illustrative diagram showing an example of the processing performed by the image analysis unit. [Figure 8] This is an illustrative diagram illustrating an example of how the image analysis unit's processing can be supplemented. [Figure 9] This is an explanatory diagram of an example of the AI-ASD algorithm. [Figure 10] This is an illustrative diagram of an example of additional processing 1. [Figure 11] This is an illustrative diagram of an example of additional processing 2. [Figure 12] This is an explanatory diagram illustrating an example of the correspondence between meeting participant names and text data. [Figure 13] This is a sequence diagram of an example of the procedure for a person identification method performed by the person identification device according to this embodiment. [Figure 14] This is an explanatory diagram illustrating an example of the content of a speech notification transmitted from the sound detection unit to the image analysis unit. [Figure 15] This is a sequence diagram of an example of the procedure for creating meeting minutes performed by the meeting minutes creation system according to this embodiment. [Figure 16] This is a diagram illustrating an example of how this embodiment is applied to a monitoring system. [Modes for carrying out the invention]
[0008] Embodiments of the present invention will be described below with reference to the attached drawings.
[0009] <System Configuration> Figure 1 is a configuration diagram of an example of a meeting minutes creation system 1 according to this embodiment. The meeting minutes creation system 1 according to this embodiment includes a person identification device 10 and a meeting minutes creation device 12. The person identification device 10 and the meeting minutes creation device 12 are connected in a communicative manner. The person identification device 10 and the meeting minutes creation device 12 are connected in a communicative manner, for example, via a data transfer cable. The person identification device 10 and the meeting minutes creation device 12 may be connected via a network such as the Internet or a LAN (Local Area Network).
[0010] The person identification device 10 identifies the speaker (the person who made the sound, which is an example of a sound produced) using audio information acquired (recorded) by the microphone and image information acquired (captured) by the camera. The person identification device 10 may use a built-in microphone and camera, or it may use an external microphone and camera. The person identification device 10 is installed in a place where audio information and image information are acquired (for example, a conference room). The person identification device 10 may be a PC (Personal Computer), smartphone, tablet terminal, mobile phone, wearable device, HMD (Head Mounted Display), imaging device, or dedicated equipment.
[0011] The minutes creation device 12 converts audio information into text data (character data). The minutes creation device 12 creates minutes by associating the speaker identified by the person identification device 10 with the text data obtained by converting the audio information. The minutes creation device 12 is, for example, a PC, a workstation, or the like. The minutes creation device 12 may be a smartphone, a tablet terminal, a mobile phone, a wearable terminal, or the like. The minutes creation device 12 may be a printer, a scanner, a facsimile machine, a multifunction peripheral, a projector, a display device having an electronic blackboard function, or the like.
[0012] Note that the configuration of the minutes creation system 1 shown in FIG. 1 is an example. The configuration of the minutes creation system 1 varies depending on the application, purpose, and the like. For example, the person identification device 10 and the minutes creation device 12 may be configured by a plurality of computers. Further, the minutes creation device 12 may be implemented as a cloud computing service. The person identification device 10 and the minutes creation device 12 may have an integrated configuration. Further, the function of the person identification device 10 may be provided as one function of the minutes creation device 12.
[0013] <Hardware Configuration> FIG. 2 is a hardware configuration diagram of an example of the person identification device 10 according to the present embodiment.
[0014] The person identification device 10 includes a CPU (Central Processing Unit) 401, a ROM (Read Only Memory) 402, a RAM (Random Access Memory) 403, an EEPROM 404, and a media I / F (Interface) 409.
[0015] The CPU 401 is an example of a processor that controls the operation of the person identification device 10. The processor that controls the operation of the person identification device 10 may include a device such as a GPU (Graphics Processing Unit).
[0016] The ROM 402 stores programs used for driving the CPU 401, such as the CPU 401 and an IPL. The RAM 403 is used as a work area for the CPU 401. The EEPROM 404 reads or writes various data such as a program for the person identifying device 10 under the control of the CPU 401. The media I / F 409 controls reading or writing (storage) of data to / from a recording medium 408 such as a flash memory.
[0017] Further, the person identifying device 10 includes a camera 413, an image sensor I / F 414, a microphone 415, a speaker 416, an audio input / output I / F 417, a display 418, an external device connection I / F 419, a near field communication circuit 420, an antenna 420a for the near field communication circuit 420, and a touch panel 421.
[0018] The camera 413 is a type of built-in imaging means that acquires (captures) image information in accordance with control by the CPU 401. The image sensor I / F 414 is a circuit that controls driving of the camera 413. The microphone 415 is a type of built-in recording means that converts sound into electrical signals and acquires (records) audio information. The speaker 416 is a built-in circuit that converts electrical signals into physical vibrations to generate sound.
[0019] The audio input / output I / F 417 is a circuit that processes input / output of audio information between the microphone 415 and the speaker 416 in accordance with control by the CPU 401. The display 418 is a type of display means such as a liquid crystal display or an organic EL (Electro Luminescence) display.
[0020] The external device connection interface 419 is an interface for connecting various external devices. The near-field communication circuit 420 is a communication circuit such as NFC (Near Field Communication) or Bluetooth (registered trademark). The touch panel 421 is a type of input means that allows the operator to operate the person identification device 10 by pressing the display 418. The person identification device 10 is equipped with a bus line 410. The bus line 410 is an address bus and data bus, etc., for electrically connecting each component such as the CPU 401 shown in Figure 2.
[0021] Note that the hardware configuration shown in Figure 2 is just one example, and it is not necessary to include all of the components shown in Figure 2, or to include components other than those shown in Figure 2.
[0022] The meeting minutes creation device 12 shown in Figure 1 may also be implemented using a computer 500 with the hardware configuration shown in Figure 3. Figure 3 is a hardware configuration diagram of an example of computer 500.
[0023] Computer 500 is equipped with a CPU 501, ROM 502, RAM 503, HD 504, HDD (Hard Disk Drive) controller 505, display 506, external device connection I / F 508, network I / F 509, data bus 510, keyboard 511, pointing device 512, DVD-RW (Digital Versatile Disk Rewritable) drive 514, and media I / F 516.
[0024] The CPU 501 controls the operation of the entire computer 500 according to the program. The ROM 502 stores programs used to drive the CPU 501, such as the IPL. The RAM 503 is used as the work area for the CPU 501. The HD 504 stores various data, such as programs. The HDD controller 505 controls the reading or writing of various data to the HD 504 according to the control of the CPU 501.
[0025] The display 506 displays various information such as cursors, menus, windows, characters, or images. The external device connection interface 508 is an interface for connecting various external devices. In this case, external devices include, for example, USB (Universal Serial Bus) memory. The network interface 509 is an interface for data communication using the network 18. The data bus 510 is an address bus and data bus for electrically connecting various components such as the CPU 501.
[0026] The keyboard 511 is a type of input means equipped with multiple keys for inputting characters, numbers, and various instructions. The pointing device 512 is a type of input means for selecting and executing various instructions, selecting processing targets, and moving the cursor. The DVD-RW drive 514 controls the reading or writing of various data to the DVD-RW 513, which is an example of a removable recording medium. Note that it is not limited to DVD-RW, but may also be DVD-R, etc. The media interface 516 controls the reading or writing (storage) of data to the recording medium 515, such as flash memory.
[0027] Note that the hardware configuration shown in Figure 3 is just one example, and it is not necessary to include all of the components shown in Figure 3, or to include components other than those shown in Figure 3.
[0028] <Functional Configuration> Figure 4 is a functional configuration diagram of an example of the meeting minutes creation system 1 according to this embodiment. In the functional configuration diagram of Figure 4, components unnecessary for the explanation of this embodiment have been appropriately omitted. The person identification device 10 realizes the functional configuration of Figure 4 by executing the OS (Operating System) and program with the hardware configuration shown in Figure 2, for example. Similarly, the meeting minutes creation device 12 realizes the functional configuration of Figure 4 by executing the OS and program with the hardware configuration shown in Figure 3, for example.
[0029] The person identification device 10 shown in Figure 4 includes a microphone unit 20, a sound detection unit 22, an image analysis unit 24, and a camera unit 26. The meeting minutes creation device 12 shown in Figure 4 includes a meeting minutes creation unit 30, a transcription unit 32, a participant management unit 34, a person tracking unit 36, and a storage unit 38.
[0030] The microphone unit 20 acquires (records) audio information. The microphone unit 20 also detects the direction of the sound using a microphone array. A microphone array is a configuration in which multiple microphones 415 are arranged on a plane, for example, using three microphones 415 to detect directions in a 360° range. In order to detect directions in a 360° range, the microphone array consists of three or more microphones 415 arranged on a plane.
[0031] For example, the signal output from microphone 415 is taken into CPU 401 via an I / F such as I2S (Inter-IC Sound) or DMIC, and processed by a DSP (Digital Signal Processor) to be converted into audio data. The audio data is sampled, for example, at 48kHz and 16-bit.
[0032] The camera unit 26 acquires (captures) image information. The camera unit 26 acquires 360-degree image information of the surroundings using a camera 413 such as a 360° panoramic camera. For example, the camera unit 26 uses a camera 413 such as a 360° panoramic camera to capture the faces of all participants in a meeting. As an example, the signal from the CMOS image sensor output from the camera 413 is taken into the CPU 401 via an I / F such as MIPI (Mobile Industry Processor Interface), and the image is processed by an ISP (Image Signal Processor) or GPU to create image data. The image format is 1080p / 30fps or 720p / 30fps, etc.
[0033] The sound detection unit 22 detects speech and the direction of speech from the audio information acquired by the microphone unit 20. For example, the sound detection unit 22 can detect speech if the audio information acquired by the microphone unit 20 is at a frequency that can be identified as a human voice (e.g., 100Hz to 4000Hz) and the sound pressure is 50dB or higher for a predetermined time of 1 second or more. The sound detection unit 22 also detects the direction of speech from the difference in arrival times of sounds detected by multiple microphones 415 of the microphone array.
[0034] The person tracking unit 36 performs person tracking on the image information output by the camera unit 26 using a person tracking algorithm, thereby managing tracking information such as when and where in the image a meeting participant appeared.
[0035] Figure 5 is an illustrative diagram of an example of the processing performed by the person tracking unit 36. In the example shown in Figure 5, the person tracking unit 36 performs person tracking on the image information output by the camera unit 26 using a person tracking algorithm, thereby managing tracking information such as when and where meeting participants A, B, and C appeared in the image.
[0036] The tracking information is used by the image analysis unit 24 to identify the speaker. For example, the tracking information is used to identify which of the meeting participants A, B, and C is the speaker. Person tracking by the person tracking unit 36 can be implemented, for example, by a person tracking algorithm that combines YOLO (You Only Look Once) and ByteTrack (see, for example, https: / / github.com / ifzhang / ByteTrack). YOLO is an example of an algorithm that detects objects from an image. ByteTrack is an example of an algorithm that tracks objects detected from an image.
[0037] The participant management unit 34 associates the participants of a meeting, whose identity is being tracked by the person tracking unit 36, with pre-configured meeting participant names, and manages participant information such as when and where each meeting participant appeared in an image.
[0038] Figure 6 is an illustrative diagram of an example of the processing performed by the participant management unit 34. In the example shown in Figure 6, the person tracking unit 36 associates participants A, B, and C of the meeting, whose names are set in advance, with the names of the meeting participants (Sato, Suzuki, and Saito). The participant management unit 34 manages participant information, such as when and where in the image each meeting participant (Sato, Suzuki, and Saito) appeared.
[0039] The image analysis unit 24 analyzes the image information acquired by the camera unit 26 and identifies the speaker as follows. The image analysis unit 24 may also display the results of the speaker identification on the display 418.
[0040] First, the image analysis unit 24 determines whether the audio information acquired by the microphone unit 20 is the voice of a speaker or some other noise. Figure 7 is an illustrative diagram of an example of the processing by the image analysis unit 24. The image analysis unit 24 uses a predetermined range of ±15° from the direction of speech origin of the audio information acquired by the microphone unit 20 as the image analysis range for person detection, and applies it to a person detection algorithm to determine the presence or absence of a person. The direction of speech origin of the audio information acquired by the microphone unit 20 (the direction of the conversational audio in Figure 7) is determined when a predetermined sound pressure of 50 dB or more continues for a predetermined time of 1 second or more.
[0041] If the image analysis unit 24 determines that a person is present within the image analysis range for person detection, it identifies the speaker, as described below. If the image analysis unit 24 determines that there is no person within the image analysis range for person detection, it determines that the audio information acquired by the microphone unit 20 is a sound other than the speaker's voice. When determining whether or not a person is present within the image analysis range for person detection, it is preferable to align the origin 0° of the direction of speech generation (speech direction) of the audio information acquired by the microphone unit 20 with the 0 pixel in the horizontal direction of the image, as shown in Figure 8. Figure 8 is an illustrative diagram illustrating an example of the processing of the image analysis unit 24. In Figure 8, the range is ±15° in the horizontal direction, but a similar predetermined range (for example, ±15°) may also be used as the image analysis range in the vertical direction. Furthermore, the person detection algorithm should ideally use a model trained on face and shoulder shapes using YOLO, an example of an object detection algorithm (see, for example, https: / / ja.wikipedia.org / wiki / %E7%89%A9%E4%BD%93%E6%A4%9C%E5%87%BA).
[0042] When the image analysis unit 24 determines that there is a person within the image analysis range for person detection, it needs to identify which person is the speaker, as there may be multiple people in the direction of speech. The image analysis unit 24 identifies the speaker using an AI-ASD (Active Speaker Detection) algorithm that uses audio information and image information. At this time, by limiting the AI-ASD algorithm to image frames during a period in which the state of the audio information acquired by the microphone unit 20 was above a predetermined sound pressure for a predetermined time or longer, the image analysis unit 24 can shorten the processing time of the AI-ASD algorithm. The image analysis unit 24 can reduce the processing load of identifying the speaker from the image information acquired by the camera unit 26, thereby shortening the time required to identify the speaker.
[0043] Figure 9 is an explanatory diagram of an example of the AI-ASD algorithm. The AI-ASD algorithm uses audio and image information, and in particular, identifies speakers from 3D-recognized facial movements. The AI-ASD algorithm does not simply look at mouth movements; it can identify speakers to some extent even when they are wearing masks (see, for example, https: / / github.com / TaoRuijie / TalkNet-ASD).
[0044] To further reduce the processing time of the AI-ASD algorithm, it is advisable to reduce the image size analyzed by the AI-ASD algorithm. Therefore, either additional processing 1 or additional processing 2 below should be performed.
[0045] Figure 10 is an illustrative diagram of an example of additional processing 1. The image analysis range for person detection from the direction of speech origin of the voice information acquired by the microphone unit 20 may be changed according to the sound pressure of the speech acquired by the microphone unit 20.
[0046] Figure 10(A) shows an example where, when the sound pressure of the speech acquired by the microphone unit 20 is less than a predetermined sound pressure of 50 dB (for example, 40 dB or more but less than 50 dB), the image analysis range for person detection from the direction of speech origination of the speech acquired by the microphone unit 20 is widened from the default ±15° to ±20°.
[0047] Figure 10(B) shows an example where, when the sound pressure of the speech acquired by the microphone unit 20 is greater than a predetermined sound pressure of 50 dB (for example, 60 dB or more), the image analysis range for person detection from the direction of speech origination of the speech acquired by the microphone unit 20 is narrowed from the default ±15° to ±10°.
[0048] Figure 11 is an illustrative diagram of an example of additional processing 2. The microphone unit 20 acquires audio information that has a predetermined sound pressure and lasts for a predetermined period of time or longer. Face recognition is then performed on a specific image frame from among multiple image frames included in that period (the period for analysis). Then, the face region containing the face is detected from the face position information obtained from the specific image frame through face recognition, and the face region is extracted.
[0049] The image analysis unit 24 associates speaker information (speaker identification information) identified by an AI-ASD algorithm using audio and image information with meeting participant information. The speaker identification information includes, for example, the time coordinates of meeting participants A, B, and C. The participant information includes the time coordinates of meeting participants (Sato, Suzuki, and Saito). The image analysis unit 24 can identify the speaker by selecting the time coordinates of the identified speaker from the time coordinates of the meeting participants (Sato, Suzuki, and Saito).
[0050] The transcription unit 32 converts the audio information into text data. The transcription unit 32 may be located in a location other than the meeting minutes creation device 12; for example, a cloud service may be used. The meeting minutes creation unit 30 transmits the audio information to the transcription unit 32, causing the transcription unit 32 to perform the transcription. The meeting minutes creation unit 30 receives the text data that the transcription unit 32 has produced from the audio information.
[0051] The minutes creation unit 30 creates meeting minutes by associating the names of meeting participants identified by the image analysis unit 24 with the text data obtained by transcribing the audio information. For example, the minutes creation unit 30 refers to the speaker identification information received from the image analysis unit 24 to identify the meeting participants who were speaking at the time of the text data, and associates the names of the meeting participants with the text data.
[0052] The minutes creation unit 30 associates the names of meeting participants with text data, for example, as shown in Figure 12. Figure 12 is an explanatory diagram of an example of the associated meeting participant names and text data. As shown in Figure 12, the minutes creation unit 30 can create meeting minutes by associating the text data obtained by transcribing audio information with the names of the meeting participants who spoke the content of that text data. The minutes creation unit 30 stores the created electronic data of the meeting minutes in the storage unit 38. When the minutes creation unit 30 receives a display request from the operator, it reads the electronic data of the meeting minutes from the storage unit 38 and displays it on the display 506 or the like.
[0053] <Processing> The person identification device 10 according to this embodiment performs processing according to the procedure shown in Figure 13, for example. Figure 13 is a sequence diagram of an example of the procedure for person identification performed by the person identification device 10 according to this embodiment.
[0054] In step S1, the microphone unit 20 transmits the audio data from the three microphones 415 as audio information to the sound detection unit 22. In step S2, the sound detection unit 22 monitors the audio information received from the microphone unit 20, measures the time difference between the audio data arriving from the three microphones 415, and detects the direction of the sound. The sound detection unit 22 also determines the direction of speech (speech direction) if the audio information acquired by the microphone unit 20 continues for a predetermined time or longer at a predetermined sound pressure level or higher. Subsequently, the sound detection unit 22 sends a speech notification to the image analysis unit 24, for example, the content shown in Figure 14. The sound detection unit 22 may also determine whether the audio information acquired by the microphone unit 20 is within a frequency range that can be identified as a human voice (e.g., 100Hz to 4000Hz), and if it is within a frequency range that can be identified as a human voice, it may send a speech notification to the image analysis unit 24, for example, the content shown in Figure 14.
[0055] Figure 14 is an explanatory diagram of an example of the content of a speech notification transmitted from the sound detection unit 22 to the image analysis unit 24. The content of the speech notification in Figure 14 includes the speech duration and the speech direction. The speech duration included in the content of the speech notification is represented by the time when the audio data exceeds a predetermined sound pressure and the time when the sound pressure exceeds the predetermined sound pressure for a predetermined period of time. The speech direction included in the content of the speech notification is the speech direction calculated from the time difference of the audio detected by the three microphones 415.
[0056] Upon receiving a speech notification from the sound detection unit 22, the image analysis unit 24 performs the image analysis in step S3. In the image analysis in step S3, the image analysis unit 24 identifies the speaker using the aforementioned AI-ASD algorithm and associates it with the meeting participant information.
[0057] Figure 15 is a sequence diagram of an example of the procedure for creating meeting minutes performed by the meeting minutes creation system 1 according to this embodiment. The processes in steps S11 to S17 are repeated from the start to the end of the meeting.
[0058] In step S11, the minutes creation unit 30 receives audio information, including audio data, from the microphone unit 20. In step S12, the minutes creation unit 30 transfers the audio data to the transcription unit 32. In step S13, the transcription unit 32 converts the audio data into text data. In step S14, the minutes creation unit 30 receives the text data corresponding to the audio data transferred to the transcription unit 32.
[0059] In step S15, the minutes creation unit 30 knows the duration (start time and end time) of the original audio data that has been converted into text data by the transcription unit 32. The minutes creation unit 30 specifies the duration of the original audio data that has been converted into text data by the transcription unit 32 and queries the image analysis unit 24 for the speaker's name.
[0060] In step S16, the image analysis unit 24, in response to an inquiry from the minutes creation unit 30 regarding the speaker's name, notifies the minutes creation unit 30 of the speaker's name corresponding to the specified period of audio data. In step S17, the minutes creation unit 30 associates the text data received in step S14 with the speaker's name notified in step S16 and creates the minutes.
[0061] As described above, the person identification device 10 according to this embodiment reduces the processing load by limiting the images to be analyzed in the process of identifying a speaker using audio information acquired by the microphone unit 20 and image information acquired by the camera unit 26. Therefore, the person identification device 10 according to this embodiment can reduce the processing time until the speaker is identified and can identify the speaker's name in a short time. For example, the person identification device 10 according to this embodiment may be a real-time person identification device that identifies the speaker's name in real time.
[0062] Furthermore, the meeting minutes creation system 1 according to this embodiment can, for example, quickly identify which person is speaking among multiple meeting participants in a conference room, and associate the identified speaker's name with the text data obtained by converting the audio data into text, thereby enabling the creation of more accurate meeting minutes in a short time. For example, the meeting minutes creation system 1 according to this embodiment may be a real-time meeting minutes creation system that creates meeting minutes during the meeting.
[0063] [Other embodiments] In this embodiment, an example of identifying speakers in a meeting has been described, but the invention is not limited to identifying speakers in a meeting. For example, the person identification device 10 according to this embodiment can also be applied to systems such as surveillance systems that identify people from audio information acquired by the microphone unit and image information acquired by the camera unit of a surveillance camera. Figure 16 is a configuration diagram of an example in which this embodiment is applied to a monitoring system 100. The surveillance system 100 is configured to include a person identification device 10a and a surveillance device 120. The person identification device 10a has the same configuration as in Figure 2, and the surveillance device 120 has the same configuration as in Figure 3. In this embodiment as well, from the image information acquired by the camera unit 26 of the person identification device 10a, an image area within a predetermined range of ±15° is set as the image analysis range for person detection based on the direction of sound generation collected by the microphone unit 20, and the presence or absence of a person is determined by applying the person detection algorithm. The direction of sound generation of the audio information acquired by the microphone unit 20 is determined when a state of sound pressure exceeding a predetermined level continues for a predetermined time or longer. In this case, the sound pressure and time required to determine the direction of generation may be changed depending on the environment in which the monitoring system 100 is installed. Furthermore, in the person identification device 10a, if the person is pre-registered, the device will identify the individual, but if the person is not registered, it may only determine whether or not it is a person.
[0064] Each of the functions of the embodiments described above can be realized by one or more processing circuits. Hereinafter, "processing circuit" as used herein includes processors programmed to execute each function by software, such as processors implemented by electronic circuits, as well as devices such as ASICs (Application Specific Integrated Circuits), DSPs (digital signal processors), FPGAs (field programmable gate arrays), and conventional circuit modules designed to execute each of the functions described above.
[0065] The apparatus described in the examples represents only one of several computing environments for carrying out the embodiments disclosed herein. The present invention is not limited by these embodiments, and the components in these embodiments include those readily conceivable to those skilled in the art, those substantially identical, and those within the scope of so-called equivalents. Furthermore, various omissions, substitutions, modifications, and combinations of components can be made without departing from the spirit of these embodiments. [Explanation of symbols]
[0066] 1. Meeting Minutes Creation System 10 Person identification device 12. Meeting Minutes Creation Device 20 Microphone section 22 Sound generation detection unit 24 Image Analysis Department 26 Camera Section 30 Minutes Preparation Department 413 Camera 415 Mike [Prior art documents] [Patent Documents]
[0067] [Patent Document 1] Japanese Patent Publication No. 2021-131595
Claims
1. A sound detection unit that detects the generated sound and the direction from which the generated sound is generated from the audio information acquired by the microphone, The image analysis department analyzes the image information acquired by the camera, It has, When the sound detection unit detects the sound and its direction of origin, the image analysis unit analyzes the image information acquired by the camera within a predetermined range from the direction of origin, using this range as the analysis area, to identify the person emitting the sound. Person identification device.
2. The aforementioned camera acquires continuous video as image information, The image analysis unit analyzes image information for a period of time or longer during which the generated sound maintains a sound pressure level above a predetermined level to identify the person emitting the sound. The person identification device according to claim 1.
3. The aforementioned camera acquires image information of the surrounding 360 degrees. The image analysis unit determines the analysis range based on the direction of occurrence from the surrounding 360-degree image information. The person identification device according to claim 2.
4. The image analysis unit changes the predetermined range according to the sound pressure of the generated sound detected by the generated sound detection unit. The person identification device according to claim 3.
5. The image analysis unit detects a face region containing a face from a specific image frame among a plurality of image frames included in the period during which the analysis is performed, and determines the face region. A person identification device according to any one of claims 1 to 4.
6. A sound detection unit that detects speech and the direction of speech generation from audio information acquired by a microphone, The image analysis department analyzes the image information acquired by the camera, The minutes preparation department, which prepares the minutes of the meeting, It has, When the sound detection unit detects the utterance and the direction of utterance, the image analysis unit analyzes the image information acquired by the camera within a predetermined range from the direction of utterance to identify the speaker. The minutes creation unit associates the speaker identified by the image analysis unit with the text data obtained by converting the audio information, and creates the minutes. Meeting minutes creation system.
7. A method of identifying a person performed by a person identification device, A sound detection step that detects the generated sound and the direction from which the generated sound is generated from the audio information acquired by the microphone, The image analysis step involves analyzing the image information acquired by the camera, It has, The image analysis step, upon detecting the sound and its direction of origin in the sound detection step, analyzes the image information acquired by the camera within a predetermined range from the direction of origin to identify the person emitting the sound. How to identify a person.
8. On the computer, A sound detection procedure that detects the generated sound and the direction from which the sound originates from audio information acquired by a microphone. Image analysis procedure for analyzing image information acquired by a camera. Make it run, The image analysis procedure, upon detecting the sound and its direction of origin in the sound detection procedure, analyzes the image information acquired by the camera within a predetermined range from the direction of origin to identify the person emitting the sound. program.
Citation Information
Patent Citations
Minutes creation support system and minutes creation support device
JP2021131595A