Data processing apparatus and data processing method

The data processing device correlates speech and video data using upper body movements to identify speakers, overcoming face concealment issues and enhancing medication history creation efficiency.

JP2025127960APending Publication Date: 2025-09-02HITACHI SOLUTIONS TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024024983
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-21
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Existing speaker identification technologies struggle to accurately identify speakers when their faces are partially or entirely hidden, such as when wearing masks, mufflers, or scarves, which complicates the creation of medication history records in pharmacies.

Method used

A data processing device and method that utilize acoustic data to identify speakers by correlating speech data with video data, linking speech start times to action start times based on upper body movements, even when faces are hidden.

Benefits of technology

Enables accurate speaker identification in environments where faces are obscured, facilitating efficient medication history generation and reducing pharmacist workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025127960000001_ABST
    Figure 2025127960000001_ABST
Patent Text Reader

Abstract

To provide a data processing apparatus capable of identifying a speaker even if the face is hidden.SOLUTION: An acoustic data processing unit 216 identifies, for each speech, based on acoustic data indicating speeches of a plurality of persons, speech data indicating a speech, a speaker who spoke, and a speech start time at which the speech is started. A person image processing unit 217 identifies a motion start time at which each motion of each person is started, from video data obtained by imaging the plurality of persons, and associates, based on the speech start time and the motion start time, an image person appearing in the video data with speech data indicating a speech made by the image person as a speaker.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a data processing device and a data processing method. [Background technology]

[0002] At dispensing pharmacies, pharmacists record a drug history (hereinafter sometimes referred to as a medication history) after dispensing medication to patients, but creating a medication history record is a burden for pharmacists, especially at dispensing pharmacies with a large number of medication administration cases (pharmacists providing medication instructions to patients when handing over prescription medication). Automatically generating a medication history from the voice of the pharmacist when dispensing medication can reduce the burden on pharmacists and verify any missed explanations at the time of medication administration.

[0003] Technology for converting speech into text data already exists. However, because dispensing tables at pharmacies typically have adjacent dispensing areas to accommodate multiple patients simultaneously, speech from adjacent dispensing areas can be heard. For this reason, a mechanism for distinguishing speakers is needed to create a medication history from speech. To distinguish speakers, it is generally necessary to learn the speech features of each speaker in advance, but it is difficult for pharmacies to learn the speech features of patients in advance.

[0004] In response to this, Patent Document 1 discloses a technology that identifies the timing of opening and closing a person's mouth from a facial area image showing the person's face, estimates the speech period during which the person is speaking, and identifies the speaker based on the speech period.

[0005] Furthermore, Patent Document 2 discloses a technique for estimating a speech period from the entire face of a target person, including the mouth, and identifying the speaker based on the speech period. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Application Publication No. 2019-8134 [Patent Document 2] Japanese Patent Publication No. 2022-63080 Summary of the Invention [Problem to be solved by the invention]

[0007] The techniques described in Patent Documents 1 and 2 have a problem in that it is difficult to identify the speaker if the face is partially or entirely hidden. For example, the technique described in Patent Document 1 requires detecting the movement of the mouth, making it difficult to identify the speaker if the person is wearing a mask. In addition, the technique described in Patent Document 2 estimates the speech period from the entire face, making it possible to identify the speaker even if the person is not wearing a mask. However, it becomes difficult to identify the speaker if the person is wearing a muffler or scarf around their neck.

[0008] An object of the present disclosure is to provide a data processing device and a data processing method that are capable of identifying a speaker even when the speaker's face is hidden. [Means for solving the problem]

[0009] A data processing device according to one aspect of the present disclosure includes an acoustic data processing unit that, based on acoustic data indicating utterances made by multiple people, identifies, for each utterance, speech data indicating the utterance, the speaker who made the utterance, and the speech start time at which the utterance began; and a person image processing unit that, based on video data depicting the multiple people, identifies, for each action of each person, the action start time at which the action began, and links, based on the speech start time and the action start time, an image person who is a person depicted in the video data, with the speech data indicating the utterance made by the image person as the speaker. [Effects of the Invention]

[0010] According to the present invention, it is possible to identify a speaker even if the speaker's face is hidden. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram illustrating an utterance information extraction system according to an embodiment of the present disclosure. [Figure 2] FIG. 1 illustrates an example of the configuration of a voice processing device. [Figure 3] 10 is a flowchart illustrating the overall process. [Figure 4] 10 is a flowchart illustrating an example of a medication instruction list generation process. [Figure 5] 10 is a flowchart illustrating an example of a target person identification process. [Figure 6] FIG. 10 is a diagram illustrating an example of a speaker identification process. [Figure 7] FIG. 10 is a diagram for explaining an example of acoustic data processing. [Figure 8] FIG. 10 is a diagram illustrating an example of an utterance included in acoustic data. [Figure 9] FIG. 10 is a diagram illustrating an example of a speech data acquisition process. [Figure 10] FIG. 10 is a diagram illustrating an example of utterance data. [Figure 11] 10 is a flowchart illustrating an example of person image processing. [Figure 12] 10 is a flowchart illustrating an example of a gesture detection process. [Figure 13] 10 is a flowchart illustrating an example of an utterance identification process. [Figure 14] FIG. 10 is an image diagram schematically illustrating the relationship between the gestures and speech of an image person. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.

[0013] FIG. 1 is a diagram illustrating an utterance information extraction system according to an embodiment of the present disclosure. The utterance information extraction system illustrated in FIG. 1 is installed in a dispensing pharmacy (store). A dispensing table 100 is installed in the pharmacy for administering medication (medication instructions when a pharmacist hands over a prescribed drug to a patient), and a pharmacist 102 and a patient 103 face each other across the dispensing table 100, and the pharmacist 102 administers medication to the patient 103. There may be multiple pairs of pharmacists 102 and patients 103, and the dispensing table 100 may be divided into multiple booths (medication areas) so that each pair can administer medication simultaneously. In addition to the pharmacist 102, a supervising pharmacist 101 is employed at the store. Note that the supervising pharmacist 101 may also serve as the pharmacist 102.

[0014] The speech information extraction system includes a camera 104 , a microphone 105 , one or more audio processing devices 108 , and a monitor device 109 .

[0015] The camera 104 captures images of the pharmacist 102 and the patient 103 when the pharmacist administers the medication by capturing images of the booth where the pharmacist 102 administers the medication to the patient 103 and the surrounding area, and transmits the image data (more specifically, video image data) acquired by the image capture to the audio processing device 108. The camera 104 is installed so that the pharmacist 102 and the patient 103 in each booth can be distinguished.

[0016] There may be multiple cameras 104. Furthermore, the installation location of the cameras 104 is not particularly limited, and for example, the cameras 104 may be installed on the dispensing table 100, or may be installed on the ceiling or wall. Furthermore, the cameras 104 may be wired cameras connected to the sound processing device 108 via a wired network 106, or may be wireless cameras connected to the sound processing device 108 via a wireless router 107.

[0017] The microphone 105 acquires voices in the vicinity of the medication area where the pharmacist 102 administers medication to the patient 103, and transmits the acquired acoustic data to the voice processing device 108. There may be a plurality of microphones 105. The location where the microphone 105 is installed is not particularly limited, and the microphone 105 may be installed on the medication table 100, for example. The microphone 105 may be a wired microphone connected to the voice processing device 108 via a wired network 106, or a wireless microphone connected to the voice processing device 108 via a wireless router 107.

[0018] The camera 104 and microphone 105 may be provided in a mobile terminal such as a tablet PC 110.

[0019] The voice processing device 108 is a data processing device that analyzes the content of speech of people involved in medication (the pharmacist 102 and the patient 103) based on image data and audio data from the camera 104 and microphone 105. In this embodiment, the image data includes images of multiple people, and the audio data includes speech made by multiple people.

[0020] The monitor device 109 displays various information.

[0021] Fig. 2 is a diagram showing an example of the configuration of the audio processing device 108. The audio processing device 108 shown in Fig. 2 has a main memory device 201, a central processing unit (CPU) 202, a storage 203, an input / output interface 204, an audio interface 205, a network interface 206, a real-time clock 207, and a video display adapter 208, and each device is connected to each other via an internal bus 209.

[0022] The main memory device 201 stores programs that define the operation of the central processing unit 202 and various information used and generated by the programs. In this embodiment, the following programs are stored: an OS (Operating System) 211, a BIOS (Basic Input / Output System) 212, a medication history generation program 213, and a voice-to-text conversion program 214.

[0023] The OS 211 controls the entire voice processing device 108. The BIOS 212 mediates between the OS 211 and each piece of hardware (203 to 208). The medication history generation program 213 and the voice-to-text conversion program 214 are programs that run on the OS 211. The medication history generation program 213 realizes a medication history generation unit that generates a medication history of the patient 103 based on the processing result of the voice-to-text conversion program 214. The voice-to-text conversion program 214 realizes a voice-to-text conversion unit that generates, for each person, text data indicating the content of the utterance made by that person, based on image data and audio data from the camera 104 and microphone 105.

[0024] Specifically, the speech-to-text conversion program 214 realizes a control unit 215, an acoustic data processing unit 216, and a person image processing unit 217 as a speech-to-text conversion unit. In the figure, the components (215 to 217) of the speech-to-text conversion unit are shown in the main memory device 201 for convenience.

[0025] The central processing unit 202 reads the program recorded in the main storage device 201, and executes the read program to realize each functional unit of the voice processing device 108 (a medication history generation unit, a voice-to-text conversion unit, etc.).

[0026] The storage 203 is a secondary storage device that stores various types of information. The input / output interface 204 is, for example, an interface that complies with the USB (Universal Serial Bus) standard and receives image data from the camera 104. The audio interface 205 is, for example, a microphone connection interface having an AUX (Auxiliary) terminal and receives audio data from the microphone 105. The network interface 206 inputs and outputs various types of information via the wired network 106. The real-time clock 207 acquires the current time. The video display adapter 208 is, for example, a graphics board and is connected to the monitor device 109. The current time may be acquired from an external device (not shown) that is connected to the input / output interface 204 or the network interface 206 and is capable of acquiring time information.

[0027] FIG. 3 is a flowchart for explaining an example of the overall processing of the audio processing device 108.

[0028] In the overall process, first, a medication instruction list generation process is executed to generate a medication instruction list, which is a list of medication instructions to be given to the patient 103 (step S101).

[0029] Then, based on the image data from the camera 104, the voice processing device 108 executes a target person identification process to identify a target person who is associated with the speech contained in the acoustic data from the microphone 105 from the image person who appears in the image data (step S102).

[0030] Then, the pharmacist 102 provides medication instruction to the patient 103 based on the medication instruction list generated in the medication instruction list generation process (step S103). At this time, the voice processing device 108 executes a speaker identification process for linking the utterance made in the medication instruction to the above-mentioned subject.

[0031] Fig. 4 is a flowchart for explaining an example of the medication instruction list generation process in step S101 of Fig. 3. The medication instruction list generation process is started, for example, when the pharmacist 102 and the patient 103 are face to face with the dispensing table 100 between them. The pharmacist 102 also registers his / her own personal information in the voice processing device 108 (main memory device 201 or storage 203). The personal information of the pharmacist 102 includes an ID (identification information) that identifies the pharmacist 102. The ID of the pharmacist 102 is, for example, an employee code.

[0032] In the medication instruction list generation process, the pharmacist 102 uses an input device (not shown) to register personal information of the patient 103 in the voice processing device 108 (step S201). The personal information of the patient 103 includes, for example, the name, contact information, health insurance card number, chronic illness, the presence and type of allergies, a reception number issued in advance (such as when the prescription is received), and an ID for identifying the patient 103. The patient ID may be automatically issued according to the reception number or the like.

[0033] The pharmacist 102 uses the input device to register a list of prescription drugs to be provided to the patient 103 in the voice processing device 108 (step S202). Furthermore, the pharmacist 102 uses the input device to register contraindication information related to drugs in the voice processing device 108 (step S203). Note that the contraindication information includes contraindication information related to commonly used drugs other than the drug currently prescribed.

[0034] The pharmacist 102 uses the input device to register past medication information, which is information about prescription drugs previously prescribed to the patient 103, in the voice processing device 108 (step S204). At that time, the pharmacist 102, depending on the list of prescription drugs this time, makes inquiries about doubts, etc. to the medical institution that issued the prescription, if necessary, and provides feedback to the list of prescription drugs.

[0035] Thereafter, the voice processing device 108 generates a medication instruction list based on the registered personal information, prescription list, contraindication information, and past medication information (step S205), and ends the process.

[0036] FIG. 5 is a flowchart illustrating an example of the target person identification process in step S102 of FIG.

[0037] In the process of acquiring the linked person, the pharmacist 102 uses the input device to link his / her booth with an employee code, which is his / her identification information, and registers this as prescribing pharmacist information in the voice processing device 108 (step S301). At this time, the pharmacist 102 may register the prescribing pharmacist information using fingerprint authentication or the like.

[0038] Next, the pharmacist 102 uses the input device to specify to the voice processing device 108 the reception number that has been issued in advance to the patient 103 to whom medication is to be administered (step S302).

[0039] The control unit 215 of the voice processing device 108 uses the registered prescribing pharmacist information to identify a person appearing in a specific area in the image data from the camera 104 as the pharmacist 102, and associates the ID of the pharmacist 102 as a person ID with the specific area (step S303). For example, it is predetermined that one side of the booth (here, the right side) is the pharmacist 102 and the other side (here, the left side) is the patient 103, and the control unit 215 determines the left side of the booth indicated in the prescribing pharmacist information as the specific area.

[0040] Furthermore, the control unit 215 uses the specified reception number and the personal information of the patient 103 to identify the person appearing in a specified area (for example, the left side of the booth indicated in the prescribing pharmacist information) in the image data from the camera 104 as the patient 103, and links the ID of the patient 103 to that specified area as a person ID (step S304).

[0041] Based on the processing results of steps S303 and S304, the control unit 215 generates and registers person image information linking the person ID of each person with the area in which that person appears (step S305), and ends the processing.

[0042] FIG. 6 is a diagram for explaining an example of speaker identification processing by the voice processing device 108 during medication instruction.

[0043] In the speaker identification process, the voice processing device 108 performs loop process A in which the following processes are repeated during medication instruction.

[0044] Specifically, first, the acoustic data processing unit 216 of the voice processing device 108 executes acoustic data processing to generate, for each utterance, speech data indicating the speech included in the acoustic data, based on the acoustic data from the microphone 105 (step S401). In this embodiment, the speech data is text data indicating the speech, and includes the speech start time when the speech started, the speech end time when the speech ended, and a speaker ID that identifies the speaker who made the utterance.

[0045] Next, based on the image data from the camera 104 and the speech data generated by the audio processing device 108, the person image processing unit 217 performs person image processing to link the image person appearing in the image data with the speech data indicating the speech made by the image person as a speaker (step S402).

[0046] Then, the control unit 215 updates the medication history of the image person who is a person appearing in the image data, based on the speech data linked to the image person (step S403). For example, the control unit 215 newly registers the speech data or information based on the speech data in the medication history.

[0047] FIG. 7 is a diagram for explaining an example of the acoustic data processing in step S401 of FIG.

[0048] In the acoustic data processing, the acoustic data processing unit 216 acquires acoustic data from the microphone 105 via the audio interface 205 (step S501). Fig. 8 is a diagram showing an example of speech included in the acoustic data. As shown in Fig. 8, the acoustic data includes speeches from multiple speakers A1, A2, B1, B2, etc., made in multiple booths A, B, etc.

[0049] The acoustic data processing unit 216 analyzes the acquired acoustic data and separates the speech contained in the acoustic data into multiple speaker data indicating each speaker of the speech (step S502). Note that since the technology for separating acoustic data into speaker data is known, a detailed description thereof will be omitted. Furthermore, each speaker data simply indicates the speech of a different person and does not identify the speaker. In other words, each speaker data is not linked to the image person.

[0050] The acoustic data processing unit 216 executes, for each piece of speaker data, an utterance data acquisition process for acquiring utterance data based on the speaker data (step S503), and ends the process.

[0051] FIG. 9 is a diagram for explaining an example of the speech data acquisition process in step S503 of FIG.

[0052] In the speech data acquisition process, first, the acoustic data processing unit 216 acquires the speech start time of the utterance included in the speaker data (step S601). The acoustic data processing unit 216 analyzes the target speaker data, acquires and stores features of the utterance (voice) (step S602). Note that if features have been calculated when separating the acoustic data into multiple speaker data in step S502, these features may be used.

[0053] The acoustic data processing unit 216 assigns a speaker ID to the speaker data to identify the speaker, according to the feature calculated in step S602 (step S603). For example, the acoustic data processing unit 216 compares the feature acquired this time with feature acquired in the past to determine whether the speaker of the utterance indicated by the speaker data is a new speaker. If the speaker is a new speaker, the acoustic data processing unit 216 issues a new speaker ID and assigns it to the speaker data, and if the speaker is not a new speaker, it assigns the previous speaker ID to the speaker data.

[0054] Thereafter, the acoustic data processing unit 216 determines whether or not the speaker is speaking based on the speaker data (step S604). If the speaker is speaking, the acoustic data processing unit 216 converts the utterance indicated by the speaker data into text data (step S605) and returns to the processing of step S604. On the other hand, if the speaker is not speaking, the acoustic data processing unit 216 determines that the utterance has ended and acquires the time at that time as the utterance end time (step S606).

[0055] Then, the acoustic data processing unit 216 generates data including the text data, the speech start time, the speech end time, and the speaker ID as speech data (step S607), and ends the process.

[0056] 10 is a diagram showing an example of speech data. In the example of Fig. 10, the speech data includes a microphone ID 301 that identifies the microphone 105 that acquired the speech (acoustic data), a speech start time 302, a speech end time 303, a speaker ID 304, and text data 305.

[0057] FIG. 11 is a flowchart for explaining an example of the person image processing in step S402 of FIG.

[0058] In the person image processing, the person image processing unit 217 acquires image data from the camera 104 via the input / output interface 204 (step S701).

[0059] The person image processing unit 217 divides the acquired image data into areas linked to the person IDs included in the person image information, and acquires a plurality of person image data in which the image person of the person ID appears (step S702).

[0060] The person image processing unit 217 executes a gesture detection process for each person image data to detect a gesture, which is a movement of the person depicted in the person image data (step S703).

[0061] For each person ID, the person image processing unit 217 executes a speech identification process to link the person ID of the image person with the speech made by that image person based on the gesture of the image person of that person ID and the speech data generated by the acoustic data processing unit 216 (step S704), and then terminates the process.

[0062] Fig. 12 is a flowchart for explaining an example of the gesture detection process in step S703 in Fig. 11. The gesture detection process is performed for each frame of person image data.

[0063] In the gesture detection process, the person image processing unit 217 acquires, for each frame of person image data, a feature amount of the person in the image captured in that frame (step S801). The feature amount of the person in the image is specifically a feature amount related to the upper body (more specifically, a feature amount capable of detecting a movement (gesture) of the upper body), such as the positions of predetermined parts of the human body (centers of both eyes, base of the head, shoulders, elbows, wrists, etc.), and the distance between predetermined parts.

[0064] The person image processing unit 217 calculates the amount of variation in the feature amount of the person in the image (step S802). For example, the person image processing unit 217 calculates, for each feature amount, the difference between the feature amount of the current frame and the feature amount of the frame a predetermined number (for example, 4) before as the amount of change in the feature amount.

[0065] Based on the amount of change in the feature amount, person image processing unit 217 determines whether the person in the image has made a gesture (step S803). The method for determining whether the person in the image has made a gesture is not particularly limited. For example, if there is a feature amount whose amount of change is equal to or greater than a specified value, person image processing unit 217 determines that the person has made a gesture. The specified value may be different for each feature amount.

[0066] If the person makes a gesture, person image processing unit 217 acquires the time of the current frame as the motion start time when the person made the gesture (step S804) and ends the process. On the other hand, if the person does not make a gesture, person image processing unit 217 skips the process of step S804 and ends the process.

[0067] FIG. 13 is a flowchart illustrating an example of the utterance identification process in step S704 of FIG.

[0068] In the utterance identification process, the person image processing unit 217 performs loop process B, which repeats the following process for each speaker ID for the target person ID. In other words, loop process B is performed for each combination of the same image person and speaker. In addition, in loop process B, the person image processing unit 217 performs loop process C, which repeats the following process for each utterance data of the target speaker ID.

[0069] Specifically, first, the person image processing unit 217 refers to the utterance start time included in the target utterance data and determines whether the utterance data is the first utterance data of the speaker with the target speaker ID (step S901).

[0070] In the case of the first utterance data, the person image processing unit 217 initializes the counter value Cn to 0 (step S902). After that, the person image processing unit 217 calculates the magnitude (absolute value) of the difference between the utterance start time Tut of the target utterance data and the gesture start time Tgs of the target person ID as a differential time δ (step S903).

[0071] The person image processing unit 217 determines whether or not there is a movement start time Tgs at which the difference time δ is smaller than the threshold value Tth (step S904).

[0072] If there is a motion start time Tgs at which the difference time δ is smaller than the threshold value Tth, the person image processing unit 217 increments the counter value Cn (step S905). If there is no motion start time Tgs at which the difference time δ is smaller than the threshold value Tth, the person image processing unit 217 skips the process of step S905.

[0073] Then, the person image processing unit 217 determines whether the counter value Cn is equal to or greater than a predetermined counter threshold Cth (step S906). If the counter value Cn is equal to or greater than the counter threshold Cth, the person image processing unit 217 identifies the image person of the target person ID as the speaker of the target speaker ID, links the target person ID to the target speaker ID, and links the speech data to the target person ID (step S907), and ends the process. If the counter value Cn is less than the counter threshold Cth, the person image processing unit 217 skips the process of step S907 and ends the process.

[0074] Furthermore, if it is determined in the processing of step S901 that this is not the first utterance, the person image processing unit 217 determines whether the counter value Cn is equal to or greater than the counter threshold Cth (step S908). If the counter value Cn is equal to or greater than the counter threshold Cth, the person image processing unit 217 determines that the speaker ID and the person ID have already been linked, and ends the processing. On the other hand, if the counter value Cn is equal to or less than the counter threshold Cth, the person image processing unit 217 proceeds to the processing of step S903.

[0075] 14 is an image diagram that schematically illustrates the relationship between gestures and speech of an imaged person in this embodiment. As shown in FIG. 14, this embodiment utilizes the fact that people often make some kind of gesture when speaking, and makes it possible to link the imaged person to the speaker based on the timing at which the imaged person starts making gestures and the timing at which the person starts speaking. It should be noted that there may be cases where a person speaks without making gestures, and in this embodiment, taking this into consideration, the speaker ID and person ID are linked when the counter value Cn is equal to or greater than the counter threshold Cth.

[0076] As described above, according to this embodiment, the acoustic data processing unit 216 identifies, for each utterance, speech data indicating the utterance, the speaker who made the utterance, and the speech start time at which the utterance began, based on acoustic data indicating utterances made by multiple people. The person image processing unit 217 identifies, from video data showing the multiple people, the action start time at which each action of each person began, and links the image person who is a person shown in the video data to the speech data indicating the utterance made by the image person as the speaker, based on the speech start time and the action start time. Therefore, because the person and the speech data are linked based on the person's actions, it is possible to identify the speaker even if the face is hidden.

[0077] In this embodiment, the movements include movements of at least the upper body, which makes it possible to more reliably identify the speaker even when the face is hidden.

[0078] In this embodiment, when there is an utterance in which the difference time δ, which is the magnitude of the difference between the utterance start time and the action start time, is less than the threshold Tth, the person image processing unit 217 links the utterance data representing the utterance to the image person who started the action at the action start time. In this case, it is possible to link the image person and the utterance data more accurately.

[0079] In this embodiment, when there are utterances with a difference time δ less than the threshold Tth for a combination of the same imaged person and the same speaker, the imaged person is linked to the utterance data of the speaker, which makes it possible to link the imaged person and the utterance data more accurately.

[0080] The above-described embodiments of the present disclosure are merely illustrative examples of the present disclosure, and are not intended to limit the scope of the present disclosure to these embodiments alone. Those skilled in the art may implement the present disclosure in various other forms without departing from the scope of the present disclosure.

[0081] For example, in the above embodiment, an example was described in which the utterance information extraction system was applied to a dispensing pharmacy, but the application of the utterance information extraction system is not particularly limited. For example, the utterance information extraction system may be applied to informed consent in a medical institution. [Explanation of symbols]

[0082] 100: Dispensing table 101: Supervising pharmacist 102: Pharmacist 103: Patient 104: Camera 105: Microphone 106: Wired network 107: Wireless router 108: Voice processing unit 109: Monitor device 110: Tablet PC 201: Main memory device 202: Central processing unit 203: Storage 204: Input / output interface 205: Audio interface 206: Network interface 207: Real-time clock 208: Video display adapter 209: Internal bus 211: OS 212: BIOS 213: Medication history generation program 214: Voice-to-text conversion program 215: Control unit 216: Acoustic data processing unit 217: Person image processing unit

Claims

1. an acoustic data processing unit that identifies, for each utterance, utterance data indicating the utterance, a speaker who made the utterance, and an utterance start time at which the utterance started, based on acoustic data indicating utterances made by a plurality of people; A data processing device having a person image processing unit that identifies the action start time at which each action of each person was started from video image data that shows the multiple people, and links the image person, who is a person shown in the video image data, to the speech data that indicates the speech made by the image person as the speaker based on the speech start time and the action start time.

2. The data processing device according to claim 1 , wherein the person image processing unit identifies the action start time based on a change in a feature amount relating to an upper body of the person in the image.

3. The data processing device according to claim 1, wherein, when there is an utterance in which the difference between the speech start time and the action start time is less than a threshold, the person image processing unit links the utterance data indicating the utterance to the image person who started the action at the action start time.

4. The data processing device according to claim 3, wherein the person image processing unit links the image person to the speech data of the speaker when there are a predetermined number or more of utterances in which the magnitude of the difference is less than the threshold for a combination of the same image person and the same speaker.

5. The data processing device according to claim 1 , wherein the speech data is text data.

6. A data processing method by a data processing device, comprising: based on acoustic data indicating utterances made by a plurality of persons, for each utterance, identifying utterance data indicating the utterance, a speaker who made the utterance, and an utterance start time at which the utterance started; identifying, from the video image data showing the plurality of people, a time at which each of the people's actions started; A data processing method that links an image person, who is a person depicted in the video image data, with the speech data indicating the speech made by the image person as the speaker based on the speech start time and the action start time.

Citation Information

Patent Citations

  • Sound source separation information detection device, robot, sound source separation information detection method and program

    JP2019008134A

  • Computer and voice processing method

    JP2022063080A