Conference recording apparatus, conference recording method, and conference recording program

The conference recording device enhances noise suppression and privacy by allowing users to set audio and video masking ranges, improving speech extraction accuracy and protecting participant privacy.

JP2025168909APending Publication Date: 2025-11-12PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024073768
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

Existing conference recording devices struggle with noise suppression, leading to decreased accuracy in extracting and separating individual speech components, and may inadvertently record unwanted audio or video from unintended directions.

Method used

A conference recording device with a processor that displays setting areas corresponding to the surroundings, allowing users to set ranges for audio and video masking, suppressing unwanted audio and video components based on direction, and enhancing speech extraction accuracy.

Benefits of technology

Improves the accuracy of speech identification and privacy protection by effectively masking noise and unwanted audio/video, ensuring clear recording of intended speech and maintaining participant privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025168909000001_ABST
    Figure 2025168909000001_ABST
Patent Text Reader

Abstract

To make it easy to perform settings for suppressing noise or the like contained in collected voice.SOLUTION: A conference recording apparatus receives, from a conference device equipped with a microphone that collects voice arriving from the surroundings and a camera that captures images of the surroundings, collected sound information collected by the microphone and video information captured by the camera, and displays a first setting area corresponding to directions of the surroundings together with the video information. A first range is set for the first setting area, and voice components arriving from directions within the first range in the collected sound information are suppressed.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a conference recording device, a conference recording method, and a conference recording program. [Background technology]

[0002] A conference device equipped with a camera and a microphone is used to capture video of multiple people participating in a conference, pick up the voices of the multiple people, and record them. In addition, there is known a technique for extracting the voices of each person from the collected voices (for example, Patent Documents 1 and 2). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 7233035 Publication [Patent Document 2] Special Publication No. 2023-546890 Summary of the Invention [Problem to be solved by the invention]

[0004] When the collected audio contains noise, the accuracy of extracting (or separating) the speech of each person from the collected audio decreases. Therefore, it is necessary to suppress noise contained in the collected audio.

[0005] An object of the present disclosure is to provide a technology that allows easy setting to suppress noise and the like contained in collected audio. [Means for solving the problem]

[0006] One aspect of the present disclosure provides a conference recording device that includes a conference device having a microphone that collects sounds coming from the surroundings and a camera that captures images of the surroundings, and an interface that receives audio information collected by the microphone and video information captured by the camera from the conference device, wherein the processor displays a first setting area corresponding to the direction of the surroundings together with the video information, a first range is set for the first setting area, and the processor suppresses audio components in the audio information that arrive from directions within the first range.

[0007] One aspect of the present disclosure provides a conference recording device that includes: an interface that receives, from a conference device that includes a microphone that collects sounds coming from the surroundings and a camera that captures images of the surroundings, audio information collected by the microphone and video information captured by the camera; and a processor; wherein the processor displays, together with the video information, a first setting area corresponding to the direction of the surroundings and a second setting area corresponding to the direction of the surroundings; a first range is set for the first setting area; a second range is set for the second setting area; and the processor suppresses audio components in the audio information that arrive from directions that do not belong to either the first range or the second range.

[0008] One aspect of the present disclosure provides a conference recording method that receives, from a conference device equipped with a microphone that picks up sound coming from the surroundings and a camera that captures an image of the surroundings, sound information picked up by the microphone and video information captured by the camera, displays a first setting area corresponding to the direction of the surroundings together with the video information, sets a first range for the first setting area, and suppresses sound components in the sound information that arrive from a direction within the first range.

[0009] One aspect of the present disclosure provides a conference recording program that causes an information processing device to receive, from a conference device having a microphone that collects sounds coming from the surroundings and a camera that captures images of the surroundings, audio information collected by the microphone and video information captured by the camera, display a first setting area corresponding to the direction of the surroundings together with the video information, set a first range for the first setting area, and suppress audio components in the audio information that arrive from directions within the first range.

[0010] One aspect of the present disclosure provides a conference recording method that receives, from a conference device equipped with a microphone that collects sound coming from the surroundings and a camera that captures images of the surroundings, sound information collected by the microphone and video information captured by the camera, displays a first setting area corresponding to the direction of the surroundings and a second setting area corresponding to the direction of the surroundings together with the video information, sets a first range for the first setting area, sets a second range for the second setting area, and suppresses sound components in the sound information that arrive from directions that do not belong to either the first range or the second range.

[0011] One aspect of the present disclosure provides a conference recording program that causes an information processing device to receive, from a conference device having a microphone that collects sound coming from the surroundings and a camera that captures images of the surroundings, sound information collected by the microphone and video information captured by the camera, displaying a first setting area corresponding to the direction of the surroundings and a second setting area corresponding to the direction of the surroundings together with the video information, setting a first range for the first setting area and a second range for the second setting area, and suppressing sound components in the sound collection information that arrive from directions that do not belong to either the first range or the second range.

[0012] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a recording medium, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium. [Effects of the Invention]

[0013] According to the present disclosure, it is possible to easily perform settings for suppressing noise and the like contained in collected audio. [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 1 is a block diagram showing a configuration example of a conference recording system according to a first embodiment. [Figure 2] FIG. 1 is a diagram for explaining a use case of the conference recording system according to the first embodiment. [Figure 3] FIG. 10 is a diagram showing an example of a setting screen according to the first embodiment; [Figure 4] FIG. 10 is a diagram showing an example of displaying warning information when people are too close to each other according to the first embodiment; [Figure 5] FIG. 10 is a diagram showing an example of a meeting recording screen according to the first embodiment. [Figure 6] FIG. 10 is a diagram for explaining a use case of the conference recording system according to the second embodiment. [Figure 7] FIG. 10 is a diagram showing an example of a setting screen according to the second embodiment; [Figure 8] FIG. 10 is a diagram showing an example of a meeting recording screen according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, embodiments of the present disclosure will be described in detail with appropriate reference to the drawings. However, more detailed description than necessary may be omitted. For example, detailed descriptions of well-known matters and redundant descriptions of substantially identical configurations may be omitted. This is to avoid unnecessary redundancy in the following description and to facilitate understanding by those skilled in the art. Note that the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure, and are not intended to limit the subject matter described in the claims.

[0016] (Embodiment 1) <Configuration> FIG. 1 is a block diagram showing an example of the configuration of a conference recording system 1 according to the first embodiment.

[0017] The conference recording system 1 includes a conference device 10 and a conference recording apparatus 20 .

[0018] The conferencing device 10 includes multiple cameras 11 capable of capturing 360-degree images in a planar view, and multiple microphones 12 capable of capturing sounds coming from all around. The conferencing device 10 may be able to detect the direction from which the sound is coming by directional analysis using the multiple microphones 12.

[0019] The conferencing device 10 generates video information based on images captured by the camera 11, and generates collected sound information based on sound collected by the microphone 12. The video information may be video data or data of multiple still images. The collected sound information may include, in addition to audio data, audio direction information indicating the direction from which the sound is coming.

[0020] The conference recording device 20 includes a processor 21, a memory 22, a storage 23, an external device I / F 24, a communication I / F 25, an input device 26, and a display device 27. I / F stands for Interface. The conference recording device 20 may be an information processing device such as a personal computer (PC), a server, a tablet terminal, or a smartphone.

[0021] The processor 21 reads and executes programs, data, etc. from the memory 22 to realize the functions of the conference recording apparatus 20. The functions of the conference recording apparatus 20 will be described as appropriate. The processor 21 may be interpreted as a central processing unit (CPU), a controller, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc. The processor may also include a graphics processing unit (GPU) and / or an artificial intelligence (AI) chip, etc.

[0022] The memory 22 stores programs, data, and the like for implementing the functions of the conference recording apparatus 20. The memory 22 may be configured by a volatile storage medium and / or a non-volatile storage medium.

[0023] The storage 23 stores programs, data, etc. for realizing the functions of the conference recording apparatus 20. The storage 23 may store video information, collected audio information, etc. received from the conference device 10. The storage 23 may store minutes information generated by the processor 21. The storage 23 may be configured with a non-volatile storage medium (e.g., a solid state drive (SSD) or a hard disk drive (HDD)).

[0024] The external device I / F 24 is connected to the conference device 10. An example of the external device I / F 24 is a USB I / F.

[0025] The communication I / F 25 is connected to an external communication network (not shown). Examples of the communication network include a wired LAN, a wireless LAN, the Internet, and a mobile communication network.

[0026] The input device 26 is a device that accepts input from a user. Examples of the input device 26 include a keyboard, a mouse, and a touch panel.

[0027] The display device 27 is a device that displays a screen. Examples of the display device 27 include a liquid crystal display, an organic EL display, and a touch panel display.

[0028] The conference recording apparatus 20 receives the video information and the collected audio information from the conference device 10 and stores them in the storage 23. The conference recording apparatus 20 also extracts (separates) the audio uttered by each person 3 (see FIG. 2 ) from the collected audio information, converts the audio of each person 3 into text, generates minutes information, and stores the minutes information in the storage 23. In other words, the conference recording apparatus 20 can automatically generate minutes information in which the speech of each person 3 is recorded as text.

[0029] <Use Case> FIG. 2 is a diagram illustrating a use case of the conference recording system 1 according to the first embodiment.

[0030] For example, as shown in FIG. 2, assume that a conference device 10 is placed in the center of a table, and multiple people 3 participating in the conference are located around the table.

[0031] In this case, the conference recording apparatus 20 generates minutes information by extracting and converting the voice components uttered by each person 3 from the collected voice information received from the conference device 10 into text, and records the minutes information in the storage 23. However, if there is a lot of noise, the accuracy of extracting the voice components uttered by each person 3 from the collected voice information and the accuracy of converting the voice components into text decrease. Furthermore, if the collected voice information includes speech from a person not participating in the conference (for example, a person in an adjacent conference room), that speech will also be recorded in the minutes information. For this reason, as shown in FIG. 2, there are cases where it is desired to exclude noise, etc. coming from a certain direction from the collected voice information collected by the conference device 10 (i.e., to mask the noise, voice, etc.).

[0032] Furthermore, from the viewpoint of privacy protection or for various other reasons, there may be cases where it is not desirable to record the video of a specific person among the video information received from the conference device 10.

[0033] Therefore, in the first embodiment, a conference recording apparatus 20, a conference recording method, and a conference recording program are described that can easily set a range for masking audio, etc., for collected audio information received from the conference device 10, and a range for masking video, etc., for video information received from the conference device 10.

[0034] <Settings screen> FIG. 3 is a diagram showing an example of the setting screen 100 according to the first embodiment.

[0035] Before starting recording of a conference, the processor 21 of the conference recording apparatus 20 displays a setting screen 100 shown in Fig. 3 on the display device 27. The display of the setting screen 100 described below may be realized by processing of the processor 21.

[0036] The setting screen 100 has a video display area 101 , an audio mask setting area 102 , a video mask setting area 103 , and an information display area 104 .

[0037] The video display area 101 displays video information received from the conferencing device 10. If the face of person 3 is detected in the video information, face detection information 105 indicating the detected position of the face may be displayed in the video display area 101.

[0038] The horizontal direction of the video display area 101 corresponds to the directions of 360 degrees around the conference device 10. For example, angle information 106A indicating a direction of 0 degrees around the conference device 10 and angle information 106B indicating a direction of 180 degrees around the conference device 10 may be displayed in the horizontal direction of the video display area 101. This allows the user to recognize the correspondence between the video information displayed in the video display area 101 and the actual direction in which the camera 11 of the conference device 10 is capturing the image.

[0039] The central angle of the video display area 101 (180 degrees in FIG. 3) may be changeable by a user operation. For example, the user may change the central angle of the video display area 101 by dragging (swiping) the video display area 101 left or right. Also, the range of the surrounding angles displayed in the video display area 101 may be smaller than 360 degrees. For example, the range of the surrounding angles displayed in the video display area 101 may be from 45 degrees to 305 degrees.

[0040] The audio mask setting area 102 is displayed above the video display area 101, extending horizontally across the video display area 101. As shown in FIG. 3, the audio mask setting area 102 is made up of a plurality of check boxes. The range in which the check boxes in the audio mask setting area 102 are selected is the range in which the audio mask is set. Hereinafter, the range in which the audio mask is set is referred to as an audio mask range 112. Note that the audio mask setting area 102 may be read as a first setting area, and the audio mask range 112 may be read as a first range.

[0041] The area 114 in which the audio mask range 112 is set in the video display area 101 may be displayed in a manner that makes it distinguishable from other areas (for example, grayed out). This allows the user to visually recognize the area 114 in which the audio mask range 112 is set in the video display area 101.

[0042] When the user selects each of the multiple check boxes in the audio mask setting area 102, the selected points may be set as the audio mask range 112. When the user selects the start and end points of the multiple check boxes in the audio mask setting area 102, the area between the start and end points may be set as the audio mask range 112. When the user drags (traces) over the multiple check boxes in the audio mask setting area 102, the dragged (traced) range may be set as the audio mask range 112.

[0043] The configuration of the audio mask setting area 102 is not limited to a plurality of check boxes. For example, the audio mask setting area 102 may be a bar extending horizontally. In this case, the user may select the start and end points of the bar in the audio mask setting area 102, thereby setting the area between the start and end points as the audio mask range 112. Furthermore, the user may drag (trace) the bar in the audio mask setting area 102, thereby setting the dragged (traced) area as the audio mask range 112.

[0044] The processor 21 of the conference recording device 20 masks the sounds in the sound collection information that come from directions within the sound mask range 112. For example, the processor 21 suppresses the sound components in the sound collection information that come from directions within the sound mask range 112. This suppresses noise that is unnecessary for recording the conference, improving the accuracy of speaker identification when creating the minutes.

[0045] The video mask setting area 103 is displayed below the video display area 101, extending horizontally across the video display area 101. The video mask setting area 103 is made up of a plurality of check boxes, as shown in FIG. 3, for example. The range in which the check boxes in the video mask setting area 103 are selected is the range in which the video mask is set. Hereinafter, the range in which the video mask is set is referred to as the video mask range 113. Note that the video mask setting area 103 may be read as a second setting area, and the video mask range 113 may be read as a second range.

[0046] In video display area 101, the area in which video mask range 113 is set may be displayed in a manner that allows it to be distinguished from other areas (not shown in FIG. 3). This allows the user to visually recognize the area in video display area 101 in which video mask range 113 is set.

[0047] When the user selects each of the multiple check boxes in video mask setting area 103, the selected points may be set as video mask range 113. Also, when the user selects the start and end points of the multiple check boxes in video mask setting area 103, the area between the start and end points may be set as video mask range 113. Also, when the user drags (traces) over the multiple check boxes in video mask setting area 103, the dragged (traced) area may be set as video mask range 113.

[0048] Like the above-described audio mask setting area 102, the video mask setting area 103 is not limited to a plurality of check boxes, and may be, for example, a bar extending horizontally. Furthermore, the audio mask setting area 102 and the video mask setting area 103 do not both need to be check boxes, and may have different configurations. Furthermore, the audio mask setting area 102 may be located below the video display area 101, and the video mask setting area 103 may be located above the video display area 101.

[0049] The processor 21 of the conference recording device 20 masks people who appear within the image mask range 113 in the image information.

[0050] For example, the processor 21 erases or does not record a portion of the video information that is within the video mask range 113. Here, erasing may mean erasing that portion after the video information has been recorded in the memory 22 or the storage 23. Furthermore, not recording may mean not recording that portion when recording the video information in the storage 23.

[0051] Alternatively, processor 21 may process the facial portion of person 3 that appears within image mask range 113 in the image information so that person 3 cannot be identified. Here, processing so that person 3 cannot be identified may involve applying masking or mosaic to the facial portion of person 3 detected in the image information.

[0052] <When people are too close> FIG. 4 is a diagram showing an example of displaying warning information when people are too close to each other according to the first embodiment.

[0053] If the distance between people is too close, the accuracy of extracting (separating) the voice components of each person 3 from the sound pickup information decreases. As a result, there is a high possibility that the correspondence between the person 3 and the speech content will be incorrect when creating the minutes. Therefore, if the distance between people is less than a predetermined threshold, the processor 21 displays warning information on the setting screen 100 indicating that the distance between the people 3 is too close.

[0054] For example, as shown in Fig. 4, processor 21 displays message 121, which is an example of warning information, saying "People displayed in frames, please keep an appropriate distance from the person next to you," in information display area 104. Also, as shown in Fig. 4, processor 21 displays frame 122, which is an example of warning information, surrounding a person who is close, in video display area 101. This allows conference participants to recognize from setting screen 100 people who are too close.

[0055] <Meeting Record> When the user clicks the voice recognition start button 107 after setting the voice mask range 112 and the video mask range 113 on the setting screen 100, the conference recording apparatus 20 records the collected voice information and video information received from the conference device 10 in the storage 23 based on these settings. Note that the user may choose whether or not to perform the recording.

[0056] Furthermore, the processor 21 of the conference recording device 20 may generate the minutes information of a conference using the following method. First, the processor 21 suppresses, from the sound collection information, sound components arriving from directions within the sound mask range 112. Next, the processor 21 identifies which person 3 spoke which sound component included in the suppressed sound collection information, based on sound direction information included in the suppressed sound collection information and the direction in which the person is located detected from the video information, using, for example, the method described in Patent Document 1. Next, the processor 21 converts the identified sound component into text to generate a speech text, and records the speech text in association with the person 3 who spoke the sound component in the minutes information.

[0057] In this way, by using sound collection information in which sound components (noise) arriving from directions within the sound mask range 112 are suppressed, the accuracy of identifying which sound component is spoken by which person 3 is improved. Also, as described above, if the distance between the people 3 is too close on the setting screen 100, the accuracy of identifying which sound component is spoken by which person 3 is further improved by increasing the distance between the people in advance.

[0058] <Recorded minutes> FIG. 5 is a diagram showing an example of the proceedings recording screen 150 according to the first embodiment.

[0059] The conference recording apparatus 20 displays the contents of the minutes information on the minutes recording screen 150. The minutes recording screen 150 may be displayed in real time during the conference, or may be displayed after the conference in response to a user operation.

[0060] For example, as shown in FIG. 5, on a proceedings recording screen 150, person information 151 and utterance text 152 are displayed in association with each other.

[0061] The person information 151 is information for identifying the person 3 participating in the conference. The person information 151 may be a captured image of the person 3 detected in the video information. The captured image of the person 3 may be an image of the face of the person 3, or an image of the body of the person 3. However, the person information 151 of the person 3 located within the video mask range 113 may not be a captured image of the person 3, but may be, for example, a predetermined icon, the name, nickname, or ID of the person.

[0062] The proceedings recording screen 150 displays person information 151 and utterance text 152 uttered by the person in association with each other. The proceedings recording screen 150 also displays the person information 151 of each person in association with the utterance text 152 in chronological order (for example, from top to bottom of the screen). This allows the user to easily check from the proceedings recording screen 150 which person 3 made what utterance.

[0063] (Embodiment 2) In the second embodiment, a use case of the conference recording system 1 different from that in the first embodiment will be described. The configuration of the conference recording system 1 in the second embodiment may be the same as that in the first embodiment.

[0064] <Use Case> FIG. 6 is a diagram illustrating a use case of the conference recording system 1 according to the second embodiment.

[0065] For example, as shown in FIG. 6, assume that the conference device 10 is placed in the center of a table, with a staff member 4 sitting on one side of the table and a customer 5 sitting on the other side of the table.

[0066] In this case, the conference recording apparatus 20 generates minutes information by converting the voice components uttered by the staff member 4 and the customer 5 from the collected voice information received from the conference device 10 into text, and records the minutes information in the storage 23. However, if there is a lot of noise, the accuracy of extracting the voice components uttered by each person from the collected voice information and the accuracy of converting the voice components into text decrease. Furthermore, if the collected voice information includes speech from a person not participating in the conference (for example, a person at an adjacent table), that speech will also be recorded in the minutes information. Therefore, as shown in FIG. 6, there are cases where it is desired to exclude from processing noise and the like coming from directions other than those from which the staff member 4 and the customer 5 are seated at the same table among the collected voice information collected by the conference device 10 (masking noise, voice, etc.).

[0067] Furthermore, from the viewpoint of privacy protection or for various other reasons, there may be cases where it is not desirable to record the video information received from the conference device 10, particularly the video of the customer 5.

[0068] Therefore, in the second embodiment, a conference recording apparatus 20, a conference recording method, and a conference recording program are described that can easily set a range for masking audio, etc. for collected audio information received from the conference device 10, and a range for masking video, etc. for video information received from the conference device 10.

[0069] The relationship between staff member 4 and customer 5 in embodiment 2 is an example of a relationship between people with different roles, and staff member 4 may be interpreted as a person with a first role (first person), and customer 5 as a person with a second role (second person). Other examples of the relationship between a person with a first role and a person with a second role include a doctor and a patient, an interrogator and a suspect, etc. Staff member 4 and customer 5 may also be interpreted simply as person 3.

[0070] <Settings screen> FIG. 7 is a diagram showing an example of the setting screen 200 according to the second embodiment.

[0071] Before starting recording of a conference, the processor 21 of the conference recording apparatus 20 displays a setting screen 200 shown in Fig. 7 on the display device 27. The display of the setting screen 200 described below may be realized by processing of the processor 21. The user may be able to select between the setting screen 100 shown in Fig. 3 and the setting screen 200 shown in Fig. 7.

[0072] The setting screen 200 has a video display area 201 , a staff setting area 202 , a customer setting area 203 , and an information display area 104 .

[0073] The video display area 201 is similar to the video display area 101 described in the first embodiment, and therefore a description thereof will be omitted here.

[0074] The staff setting area 202 is displayed above the video display area 201, extending horizontally across the video display area 201. As shown in FIG. 7, the staff setting area 202 is made up of a plurality of check boxes. The range in which the check boxes in the staff setting area 202 are selected is the range set as the range in which staff 4 is present. Hereinafter, the range set as the range in which staff 4 is present is referred to as staff range 212. Note that the staff setting area 202 may be read as a first setting area, and the staff range 212 may be read as a first range. The method for setting the staff range 212 in the staff setting area 202 may be the same as in embodiment 1.

[0075] The customer setting area 203 is displayed below the video display area 201, extending horizontally across the video display area 201. The customer setting area 203 is configured with a plurality of check boxes, as shown in FIG. 7, for example. The range in which the check boxes in the customer setting area 203 are selected is the range set as the range in which the customer 5 is present. Hereinafter, the range determined to be set by the customer 5 is referred to as the customer range 213. The customer setting area 203 may be read as the second setting area, and the customer range 213 may be read as the second range. The method for setting the customer range 213 in the customer setting area 203 may be the same as in the first embodiment.

[0076] Areas 214 in the video display area 101 where neither the staff range 212 nor the customer range 213 is set may be displayed in a manner that makes them distinguishable from other areas (for example, grayed out). This allows the user to visually recognize areas in the video display area 101 where an audio mask has been set.

[0077] Furthermore, if the distance between people is too close, the setting screen 200 may display warning information, as in the first embodiment.

[0078] The processor 21 of the conference recording device 20 may mask the customers 5 appearing within the customer area 213 in the video information.

[0079] For example, the processor 21 erases or does not record the portion of the video information that is within the customer range 213. Here, "erasing" may mean erasing that portion after the video information has been recorded in the memory 22 or the storage 23. Furthermore, "not recording" may mean not recording that portion when recording the video information in the storage 23.

[0080] Alternatively, the processor 21 may process the face of the customer 5 that appears within the customer range 213 in the video information so that the customer 5 cannot be identified. Here, processing the face of the customer 5 so that the customer 5 cannot be identified may involve applying masking or mosaic to the face of the customer 5 that is detected in the video information.

[0081] <Meeting Record> When the user clicks the voice recognition start button after setting the staff range 212 and customer range 213 on the setting screen 200, the conference recording apparatus 20 records the picked-up voice information and video information received from the conference device 10 in the storage 23 based on these settings. Note that the user may choose whether or not to perform the recording.

[0082] The processor 21 of the conference recording device 20 may generate conference minutes information by the following method. First, the processor 21 suppresses, from the sound pickup information, sound components coming from directions that do not belong to either the staff range 212 or the customer range 213. Next, from the sound pickup information after suppression, the processor 21 extracts sound components coming from directions within the staff range 212 as staff voices, and extracts sound components coming from directions within the customer range 213 as customer voices. Next, the processor 21 converts the extracted staff voices into text to generate staff utterance text, and converts the extracted customer voices into text to generate customer utterance text, and records the staff utterance text and the customer utterance text in the conference minutes information so that they can be distinguished from each other. Note that the staff utterance text may be read as a first utterance text, and the customer utterance text may be read as a second utterance text.

[0083] In this way, by setting the staff range 212 and the customer range 213 and using collected sound information in which sound components coming from directions outside the staff range 212 and the customer range 213 are suppressed, the accuracy of identifying the sound components uttered by the staff member 4 and the sound components uttered by the customer 5 is improved. Furthermore, by setting the staff range 212 and the customer range 213 in advance, the utterances of the staff member and the customer can be recorded in the minutes information so that they can be distinguished from each other.

[0084] <Recorded minutes> FIG. 8 is a diagram showing an example of a proceedings recording screen 250 according to the second embodiment.

[0085] The conference recording apparatus 20 displays the contents of the minutes information described above on the minutes recording screen 250. The minutes recording screen 250 may be displayed in real time during the conference, or may be displayed after the conference in response to a user operation.

[0086] For example, as shown in FIG. 8, on a minutes recording screen 250, staff information 251 and staff utterance text 252 are displayed in association with each other, and customer information 253 and customer utterance text 254 are displayed in association with each other.

[0087] The staff information 251 is information for identifying the staff member 4 participating in the conference. The staff member information 251 may be a photographed image of the staff member 4 detected in the video information. The photographed image of the staff member 4 may be an image of the face of the staff member 4, or an image of the body of the staff member 4. However, the staff member information 251 may also be an icon of the staff member, the name, nickname, or ID of the staff member, etc.

[0088] The customer information 253 is information indicating the customer 5 participating in the conference. The customer information 253 may not be a photographed image, but may be, for example, a predetermined icon, the customer's name, nickname, or ID. Alternatively, the customer information 253 may simply be displayed as "Customer." This allows the privacy of the customer 5 to be protected.

[0089] On the minutes recording screen 250, staff information 251 and staff utterance text 252 are displayed in association with each other, and customer information 253 and customer utterance text 254 are displayed in association with each other. For example, as shown in FIG. 8, the left side of the minutes recording screen 250 may display the staff information 251 and staff utterance text 252 in association with each other, and the right side of the minutes recording screen 250 may display the customer information 253 and customer utterance text 254 in association with each other. Furthermore, the minutes recording screen 250 displays the utterances of staff member 4 and customer 5 in chronological order (for example, from top to bottom of the screen). This allows the user to easily check from the minutes recording screen 250 what kind of conversation took place between staff member 4 and customer 5.

[0090] Summary of the Disclosure Based on the above description of the present disclosure, the following techniques are disclosed.

[0091] <Technology 1> In a conference recording device (20) comprising a conference device (10) having a microphone (12) for collecting sounds coming from the surroundings and a camera (11) for capturing images of the surroundings, an interface (e.g., external device I / F 24) for receiving sound information collected by the microphone and video information captured by the camera, and a processor (21), the processor displays a first setting area (e.g., audio mask setting area 102) corresponding to the direction of the surroundings together with the video information, sets a first range (e.g., audio mask range 112) for the first setting area, and suppresses audio components in the sound information coming from a direction within the first range. By setting the first range in the direction from which sounds unnecessary for recording the meeting, such as noise and sounds from adjacent conference rooms, are coming, the sound components unnecessary for recording the meeting are suppressed, and the sounds necessary for recording the meeting can be obtained more clearly.

[0092] <Technology 2> In the conference recording device described in Technology 1, the processor identifies the direction in which a person is located from the video information, extracts the person's voice components from the audio pickup information based on the identified direction in which the person is located, generates a speech text by converting the extracted voice components into text, and records or displays the person information corresponding to the identified person and the speech text of the identified person in association with each other. This allows the audio components spoken by a person identified from the video information to be accurately extracted from the audio pickup information, and the personal information indicating the person and the spoken text spoken by the person to be accurately associated and recorded or displayed.

[0093] <Technology 3> In the conference recording device described in Technique 2, the person information corresponding to the person is information based on an image of the person captured by the camera. This allows an image of a person to be recorded or displayed in association with the spoken text of the person, thereby making it possible to easily identify the person who spoke the spoken text.

[0094] <Technology 4> In the conference recording device described in Technology 2 or 3, the processor displays a second setting area (e.g., a video mask setting area 103) corresponding to the surrounding direction together with the video information, and a second range (e.g., a video mask range) is set for the second setting area, and the person information in the video information corresponding to the person located in a direction within the second range is information that is not based on an image of the person captured by the camera. By setting the second range in the direction of a person whose privacy needs to be protected, personal information (such as name, nickname, ID, etc.) that is not based on an image of the person can be recorded or displayed in association with the spoken text of the person. This makes it possible to protect the privacy of the person who spoke the spoken text.

[0095] <Technology 5> In the conference recording device according to Technology 4, when recording the video information, the processor does not record or erases the second range of the video information. This allows a person located in a direction within the second range in the video information to not be recorded or to be deleted, thereby protecting the privacy of that person.

[0096] <Technology 6> In the conference recording device described in Technology 4 or 5, when the processor records the video information, it processes the image of the person located in a direction within the second range in the video information so that the person cannot be identified. As a result, an image of a person in the video information who is positioned in a direction within the second range is processed so that the person cannot be identified, thereby protecting the privacy of the person.

[0097] <Technology 7> In the conference recording device according to any one of Techniques 1 to 6, the processor receives an input of the setting of the first range for the first setting area. This allows the user to easily set the first range for the first setting area.

[0098] <Technology 8> In the conference recording device according to any one of techniques 4 to 7, the processor receives an input of the setting of the second range for the second setting area. This allows the user to easily set the second range for the second setting area.

[0099] <Technology 9> In the conference recording device according to any one of the first to eighth aspects, the processor displays predetermined warning information when the distance between people identified from the video information is less than a predetermined threshold. In this way, if the distance between people is less than a predetermined threshold (i.e., the distance between people is too close), warning information is displayed and the people are asked to keep their distance, which improves the accuracy of extracting each person's voice components and makes it possible to accurately match people with the spoken text they have spoken.

[0100] <Technology 10> In the conference recording device described in Technique 9, the warning information is information that identifies a person who is displayed in the video information and whose distance is less than the threshold value. This allows, for example, participants in a meeting to easily recognize people whose distance is less than a predetermined threshold (i.e., people who are too close).

[0101] <Technology 11> A conference recording device includes a conference device (10) having a microphone (12) that picks up sounds coming from the surroundings and a camera (11) that photographs the surroundings, and an interface (e.g., external device I / F 24) that receives sound information picked up by the microphone and video information photographed by the camera, and a processor (21), wherein the processor displays, together with the video information, a first setting area (e.g., staff setting area 202) corresponding to the direction of the surroundings and a second setting area (e.g., customer setting area 203) corresponding to the direction of the surroundings, a first range (e.g., staff range 212) is set for the first setting area, and a second range (e.g., customer range 213) is set for the second setting area, and suppresses sound components in the sound information that arrive from a direction that does not belong to either the first range or the second range. By setting a first range in the direction where a person with a first role (e.g., staff member 4) is located and a second range in the direction where a person with a second role (e.g., customer 5) is located, unnecessary audio components for recording coming from directions outside of the range are suppressed, making it possible to obtain clearer audio necessary for recording the meeting.

[0102] <Technology 12> In the conference recording device described in Technology 11, the processor extracts from the sound pickup information a first sound component arriving from a direction within the first range and a second sound component arriving from a direction within the second range.

[0103] This makes it possible to obtain more clearly the first voice component of a person with a first role who is located in a direction within the first range and the second voice component of a person with a second role who is located in a direction within the second range.

[0104] <Technology 13> In the conference recording device described in Technology 12, the processor generates a first spoken text based on the first speech component, generates a second spoken text based on the second speech component, and records or displays the first spoken text and the second spoken text in a distinguishable manner. This makes it possible to accurately distinguish between the first utterance text uttered by the person playing the first role and the second utterance text uttered by the person playing the second role, and record or display them.

[0105] <Technology 14> In the conference recording device described in Technology 13, the processor identifies a first person located within the first range from the video information, identifies a second person located within the second range, records or displays the first person and the first spoken text in association with each other, and records or displays the second person and the second spoken text in association with each other. This allows a first person with a first role to be accurately matched with the first speech text spoken by the first person, and a second person with a second role to be accurately matched with the second speech text spoken by the second person and recorded or displayed.

[0106] <Technology 15> In a conference recording device described in any one of Techniques 11 to 14, the processor accepts input of the first range setting for the first setting area and input of the second range setting for the second setting area. This allows the user to easily set the first range for the first setting area and the second range for the second setting area.

[0107] <Technology 16> The conference recording method receives, from a conference device (10) equipped with a microphone (12) for picking up sounds coming from the surroundings and a camera (11) for photographing the surroundings, sound information picked up by the microphone and video information photographed by the camera, displays a first setting area (e.g., a sound mask setting area 102) corresponding to the direction of the surroundings together with the video information, sets a first range (e.g., a sound mask range 112) for the first setting area, and suppresses sound components in the sound information coming from a direction within the first range. By setting the first range in the direction from which sounds unnecessary for recording the meeting, such as noise and sounds from adjacent conference rooms, are coming, the sound components unnecessary for recording the meeting are suppressed, and the sounds necessary for recording the meeting can be obtained more clearly.

[0108] <Technology 17> The conference recording program causes an information processing device (e.g., conference recording device 20) to execute the following: receive, from a conference device (10) equipped with a microphone (12) that collects sounds coming from the surroundings and a camera (11) that captures images of the surroundings, audio information collected by the microphone and video information captured by the camera; display, together with the video information, a first setting area (e.g., audio mask setting area 102) corresponding to the direction of the surroundings; set a first range (e.g., audio mask range 112) for the first setting area; and suppress audio components in the audio information that arrive from a direction within the first range. By setting the first range in the direction from which sounds unnecessary for recording the meeting, such as noise and sounds from adjacent conference rooms, are coming, the sound components unnecessary for recording the meeting are suppressed, and the sounds necessary for recording the meeting can be obtained more clearly.

[0109] <Technology 18> The conference recording method receives, from a conference device (10) equipped with a microphone (12) that picks up sounds coming from the surroundings and a camera (11) that photographs the surroundings, audio information picked up by the microphone and video information photographed by the camera, and displays, together with the video information, a first setting area (e.g., a staff setting area) corresponding to the direction of the surroundings and a second setting area (e.g., a customer setting area) corresponding to the direction of the surroundings, a first range (e.g., a staff range 212) is set for the first setting area, and a second range (e.g., a customer range 213) is set for the second setting area, and suppresses audio components in the audio information that arrive from a direction that does not belong to either the first range or the second range. By setting a first range in the direction where a person with a first role (e.g., staff member 4) is located and a second range in the direction where a person with a second role (e.g., customer 5) is located, unnecessary audio components for recording coming from directions outside of the range are suppressed, making it possible to obtain clearer audio necessary for recording the meeting.

[0110] <Technology 19> The conference recording program causes an information processing device (e.g., conference recording device 20) to execute the following: receive, from a conference device (10) equipped with a microphone (12) that collects sounds coming from the surroundings and a camera (11) that captures images of the surroundings, audio information collected by the microphone and video information captured by the camera; display, together with the video information, a first setting area (e.g., staff setting area 202) corresponding to the direction of the surroundings and a second setting area (e.g., customer setting area 203) corresponding to the direction of the surroundings; set a first range (e.g., staff range 212) for the first setting area; set a second range (e.g., customer range 213) for the second setting area; and suppress audio components in the audio information that arrive from directions that do not belong to either the first range or the second range. By setting a first range in the direction where a person with a first role (e.g., staff member 4) is located and a second range in the direction where a person with a second role (e.g., customer 5) is located, unnecessary audio components for recording coming from directions outside of the range are suppressed, making it possible to obtain clearer audio necessary for recording the meeting.

[0111] Although the embodiments have been described above with reference to the accompanying drawings, the present disclosure is not limited to such examples. It is clear that a person skilled in the art can conceive of various modifications, alterations, substitutions, additions, deletions, and equivalents within the scope of the claims, and it is understood that these also fall within the technical scope of the present disclosure. Furthermore, the components in the above-described embodiments may be combined in any manner without departing from the spirit of the invention. [Industrial Applicability]

[0112] The technology of the present disclosure is useful for recording meetings and automatically generating minutes. [Explanation of symbols]

[0113] 1. Meeting recording system 3 people 4. Staff 5 customers 10 Conferencing Devices 11 Camera 12. Microphone 20 Meeting Recording Device 21 processors 22 Memory 23 Storage 24 External device I / F 25 Communication I / F 26 Input Devices 27 Display device 100 Settings screen 101 Video display area 102 Audio mask setting area 103 Video mask setting area 104 Information display area 105 Face detection information 106A,106B Angle information 107 Voice recognition start button 112 Audio Mask Range 113 Image mask range 114,214 areas 121 Messages 122 slots 150,250 minutes recording screen 151 Personal information 152 Spoken Text 200 Settings screen 201 Video display area 202 Staff Settings Area 203 Customer settings area 212 Staff Range 213 Customer Range 251 Staff Information 252 Staff Speech Text 253 Customer information 254 customer utterance text

Claims

1. an interface for receiving, from a conference device including a microphone for collecting sounds coming from the surroundings and a camera for capturing images of the surroundings, audio information collected by the microphone and video information captured by the camera; a processor, The processor: displaying a first setting area corresponding to the surrounding direction together with the video information; a first range is set for the first set region; suppressing sound components coming from directions within the first range in the sound collection information; Meeting recording device.

2. The processor: Identifying the direction in which the person is located from the video information; extracting a voice component of the person from the sound pickup information based on the direction in which the identified person is located; Generate a speech text by converting the extracted speech components into text; record or display personal information corresponding to the identified person and the utterance text of the identified person in association with each other; The conference recording apparatus of claim 1 .

3. The person information corresponding to the person is information based on an image of the person captured by the camera. The conference recording apparatus according to claim 2 .

4. The processor: displaying a second setting area corresponding to the surrounding direction together with the video information; a second range is set for the second set area; the person information in the video information corresponding to the person located in a direction within the second range is information that is not based on an image of the person captured by the camera; 4. The conference recording apparatus according to claim 2 or 3.

5. When recording the video information, the processor does not record or erases the second range of the video information. The conference recording apparatus according to claim 4 .

6. When recording the video information, the processor processes an image of the person located in a direction within the second range in the video information so that the person cannot be identified. The conference recording apparatus according to claim 4 .

7. the processor accepts an input of the setting of the first range for the first setting region; The conference recording apparatus of claim 1 .

8. the processor accepts an input of the second range setting for the second setting region; The conference recording apparatus according to claim 4 .

9. the processor displays predetermined warning information when the distance between the people identified from the video information is less than a predetermined threshold.

3. The conference recording apparatus according to claim 1 or 2.

10. the warning information is information that identifies a person who is displayed in the video information and whose distance is less than the threshold value; The conference recording apparatus of claim 9.

11. an interface for receiving, from a conference device including a microphone for collecting sounds coming from the surroundings and a camera for capturing images of the surroundings, audio information collected by the microphone and video information captured by the camera; a processor, The processor: displaying a first setting area corresponding to the surrounding direction and a second setting area corresponding to the surrounding direction together with the video information; a first range is set for the first set area, and a second range is set for the second set area; suppressing sound components coming from directions that do not belong to either the first range or the second range in the sound collection information; Meeting recording device.

12. The processor: extracting, from the sound collection information, a first sound component arriving from a direction within the first range and a second sound component arriving from a direction within the second range; The conference recording apparatus of claim 11.

13. The processor: generating a first spoken text based on the first speech component; generating a second spoken text based on the second speech component; recording or displaying the first spoken text and the second spoken text in a distinguishable manner; The conference recording apparatus of claim 12.

14. The processor: Identifying a first person located within the first range and a second person located within the second range from the video information; recording or displaying the first person and the first spoken text in association with each other, and recording or displaying the second person and the second spoken text in association with each other; The conference recording apparatus of claim 13.

15. The processor: accepting an input of the first range setting for the first setting area and an input of the second range setting for the second setting area; The conference recording apparatus of claim 11.

16. receiving, from a conference device including a microphone for collecting sounds coming from the surroundings and a camera for capturing images of the surroundings, audio information collected by the microphone and video information captured by the camera; displaying a first setting area corresponding to the surrounding direction together with the video information; a first range is set for the first set region; suppressing sound components coming from directions within the first range in the sound collection information; How to record a meeting.

17. receiving, from a conference device including a microphone for collecting sounds coming from the surroundings and a camera for capturing images of the surroundings, audio information collected by the microphone and video information captured by the camera; displaying a first setting area corresponding to the surrounding direction together with the video information; a first range is set for the first set region; suppressing sound components coming from directions within the first range in the sound collection information; A conference recording program that causes an information processing device to execute the above.

18. receiving, from a conference device including a microphone for collecting sounds coming from the surroundings and a camera for capturing images of the surroundings, audio information collected by the microphone and video information captured by the camera; displaying a first setting area corresponding to the surrounding direction and a second setting area corresponding to the surrounding direction together with the video information; a first range is set for the first set area, and a second range is set for the second set area; suppressing sound components coming from directions that do not belong to either the first range or the second range in the sound collection information; How to record a meeting.

19. receiving, from a conference device including a microphone for collecting sounds coming from the surroundings and a camera for capturing images of the surroundings, audio information collected by the microphone and video information captured by the camera; displaying a first setting area corresponding to the surrounding direction and a second setting area corresponding to the surrounding direction together with the video information; a first range is set for the first set area, and a second range is set for the second set area; suppressing sound components coming from directions that do not belong to either the first range or the second range in the sound collection information; A conference recording program that causes an information processing device to execute the above.

Citation Information

Patent Citations

  • Speaker identification accuracy

    JP2023546890A

  • Sound collection device, sound collection method, and program

    JP7233035B2