Audio processing device, audio processing program, and audio processing method

The audio processing device optimizes recording of microphone input and output sounds using conditional controls, addressing analysis inconveniences and data volume issues.

JP2026042547APending Publication Date: 2026-03-11OKI ELECTRIC INDUSTRY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-27
Publication Date
2026-03-11

AI Technical Summary

Technical Problem

Existing audio processing devices fail to efficiently record both microphone array input signals and processed output sounds, leading to inconvenience in defect analysis due to potential hardware or processing malfunctions, and result in excessive data volume.

Method used

An audio processing device with recording control means to selectively record microphone input signals and processed output sounds based on predetermined conditions, using parameters like recording start and stop times, maximum duration, and minimum stop intervals to manage data volume.

Benefits of technology

Efficient recording of both input and output sounds facilitates defect analysis by identifying hardware issues or processing problems, while reducing data capacity strain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026042547000001_ABST
    Figure 2026042547000001_ABST
Patent Text Reader

Abstract

To provide a sound processing device capable of efficiently recording microphone input sound and processed output sound to such an extent that no inconvenience occurs in later defect analysis. [Solution] The audio processing device of the present invention is characterized by having an audio processing means that performs predetermined audio processing on an input signal input from a microphone and outputs processed output sound, a recording means that records the input signal and the processed output sound, and a recording control means that determines whether or not to record the input signal and the processed output sound on the recording means based on predetermined conditions.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an audio processing device, an audio processing program, and an audio processing method, and can be applied to, for example, a device that stores audio data. [Background technology]

[0002] For example, Patent Document 1 discloses that area sound collection is performed using a plurality of microphone arrays, and the conversational voices collected in a first speaker area (consultant) and a second speaker area (responder) are recorded by a recording device. That is, the device of Patent Document 1 records only the processed output sound that has been subjected to voice processing (area sound collection). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 7207170 [Patent Document 2] Japanese Patent Application Laid-Open No. 2017-183902 [Patent Document 3] Japanese Patent Application Laid-Open No. 2011-070084 [Patent Document 4] Japanese Patent Publication No. 2022-032721 [Patent Document 5] Patent Publication No. 2021-033030 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the device described in Patent Document 1 above, the sound picked up by the microphone array (microphone array input signal) itself is not recorded, so if an abnormality occurs in the sound pickup processing output due to a malfunction of the microphone hardware or audio processing, even if a user or an operator who is communicating with the user at a guide terminal reports that they cannot hear the sound or that the sound has been interrupted, there is an inconvenience in terms of analyzing the sound.

[0005] From the viewpoint of defect analysis, it is desirable to have two types of data: the microphone array input signal and the processed output sound. However, simply recording both the microphone array input signal and the processed output sound in a device poses problems in terms of recording capacity.

[0006] Therefore, there is a demand for an audio processing device, an audio processing program, and an audio processing method that can efficiently record microphone input sound and processed output sound to the extent that it does not cause any inconvenience in subsequent defect analysis. [Means for solving the problem]

[0007] The first aspect of the present invention is characterized by having an audio processing means for performing predetermined audio processing on an input signal input from a microphone and outputting processed output sound, a recording means for recording the input signal and the processed output sound, and a recording control means for determining whether or not to record the input signal and the processed output sound on the recording means based on predetermined conditions.

[0008] The second audio processing program of the present invention is characterized in that it functions as an audio processing means that performs predetermined audio processing on an input signal input from a microphone and outputs processed output sound, a recording means that records the input signal and the processed output sound, and a recording control means that determines whether or not to record the input signal and the processed output sound in the recording means based on predetermined conditions.

[0009] The third aspect of the present invention is an audio processing method for use in an audio processing device, the audio processing device having an audio processing means, a recording means, and a recording control means, the audio processing means performing predetermined audio processing on an input signal input from a microphone and outputting a processed output sound, the recording means recording the input signal and the processed output sound, and the recording control means determining whether or not to record the input signal and the processed output sound in the recording means based on predetermined conditions. [Effects of the Invention]

[0010] According to the present invention, microphone input sound and processed output sound can be recorded efficiently to the extent that no inconvenience occurs in later defect analysis. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a block diagram showing a functional configuration of a voice processing device according to an embodiment. [Figure 2] 1 is a perspective view of the appearance of a sound processing device according to an embodiment; [Figure 3] 4 is a flowchart showing a recording process of the audio processing device according to the embodiment. [Figure 4] 3 is a diagram showing processing rules for each parameter related to hangover processing in the voice processing device according to the embodiment. FIG. [Figure 5] 10 is a timing chart showing an example of a hangover process in the audio processing device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0012] (A) Main embodiment Hereinafter, an embodiment of a voice processing device, a voice processing program, and a voice processing method according to the present invention will be described in detail with reference to the drawings.

[0013] (A-1) Configuration of the embodiment FIG. 2 is a perspective view of the appearance of the voice processing device 100 of this embodiment.

[0014] As shown in FIG. 2, the voice processing device 100 is a device including two microphone arrays MA (MA1, MA2), a speaker SP, and a touch panel display D. The voice processing device 100 is a device that supports voice output processing (voice output function) for outputting voice (e.g., voice guidance or the voice of a far-end talker during a call) from the speaker SP to a user, and voice collection processing (voice collection function) for collecting the user's voice using the microphone arrays MA1, MA2. The voice processing device 100 of this embodiment is also equipped with a touch panel display D that can output information to a user and accept information input from the user. For example, a device such as the user-operated terminal described in Patent Document 5 can be applied as the voice processing device 100 of this embodiment. In other words, the voice processing device 100 of this embodiment can have a hardware configuration similar to that of the user-operated terminal of Patent Document 5. Note that the touch panel display D is not an essential component in the voice processing device of this embodiment, and may be omitted.

[0015] For example, the voice processing device 100 may be configured as a device that inputs and outputs information only by voice. The use of the voice processing device 100 is not limited, but it may be configured as, for example, a transaction device (vending machine) that performs transactions such as the sale of various products. In this case, the voice processing device 100 may be configured to be able to perform transactions by receiving voice input from a user (voice collection processing) and outputting voice guidance (voice output processing).

[0016] The microphone arrays MA1 and MA2 are placed at any location (in this embodiment, the surface of the audio processing device 100 facing the user) in a space where a target area (in this embodiment, the position where the user is present) exists. The positions of the microphone arrays MA1 and MA2 relative to the target area may be anywhere as long as their directivities overlap only in the target area. Each microphone array MA is composed of two or more microphones M, and each microphone M collects an acoustic signal. In this embodiment, it is described that each microphone array MA is provided with two microphones M (M1, M2) that collect acoustic signals. In other words, each microphone array MA constitutes a 2-channel microphone array. Note that the number of microphone arrays MA is not limited to two; if there are multiple target areas, it is necessary to place a number of microphone arrays MA that can cover all of the areas.

[0017] Furthermore, although there are no limitations on the relative positions of the microphone arrays MA1, MA2 and the speaker SP, this embodiment will be described assuming that they are arranged as shown in FIG. 2. This embodiment will be described assuming that a user using the sound processing device 100 is located facing the touch panel display D. In FIG. 2, the microphone arrays MA1, MA2 are arranged on the left and right sides of the touch panel display D, respectively, as seen from the user facing the touch panel display D. Also, as shown in FIG. 2, the microphone arrays MA1, MA2 are arranged at the same vertical position (height). Also, in FIG. 2, the speaker SP is arranged on the right side as seen from the user. Note that, in the sound processing device 100 of this embodiment, the arrangement positions and orientations of the microphone arrays MA1, MA2 may be similar to those of Patent Document 5 (for example, the arrangement similar to FIG. 5 of Patent Document 5).

[0018] As described above, the audio processing device 100 uses two microphone arrays MA (MA1, MA2) to perform audio collection processing (target area audio collection processing) for collecting target area audio from a sound source in the target area, and audio output processing from the speaker SP.

[0019] FIG. 1 is a block diagram showing the functional configuration of a voice processing device 100 according to this embodiment.

[0020] The audio processing device 100 includes an audio processing unit 101 , a recording control unit 102 , and a recording unit 103 .

[0021] The voice processing device 100 may be configured entirely of hardware (for example, a dedicated chip, etc.), or may be configured partially or entirely as software (a program). The voice processing device 100 may be configured, for example, by installing a program (including the voice processing program of the embodiment) in a computer having a processor and memory in addition to input / output devices (microphone arrays MA1, MA2, a touch panel display D, a speaker SP, etc.).

[0022] The audio processing unit 101 performs predetermined audio processing on the acoustic signals collected by the microphone arrays MA1 and MA2. The audio processing performed by the audio processing unit 101 is not particularly limited, but may employ area audio collection processing described in Patent Documents 1 and 2, for example.

[0023] The sound recording control unit 102 controls the recording of the acoustic signals picked up by the microphone arrays MA1 and MA2 and the processed output sound processed by the sound processing unit 101 into the recording unit 103.

[0024] In this embodiment, if the acoustic signals and processed output sounds picked up by each microphone array MA1, MA2 were constantly recorded, the volume of recorded data would become enormous. Therefore, the recording volume is reduced by recording only the sections in which the user is speaking into each microphone array MA (each microphone M).

[0025] For example, the recording control unit 102 can define the following parameters to limit the maximum amount of recorded data per day, define the maximum amount of recorded data overall, and delete the oldest recorded data, thereby preserving recorded data for a certain period of time. This makes it possible to analyze reports made within that period.

[0026] Recording start time P1: For example, this is set to the time when the voice processing device 100 starts operating.

[0027] Recording stop time P2: For example, this is set to the time when the operation of the voice processing device 100 is to be stopped.

[0028] Maximum recording time P3: This is the maximum time from the start of recording to the end of recording. If the recording is continued for longer than this time, the recording will stop.

[0029] Minimum recording stop time length P4: This is the minimum time from when recording stops until the next recording starts, and recording will not start even if a recording start trigger (described later) is detected during this time.

[0030] For example, if the recording start time P1 = 10:00, recording stop time P2 = 20:00, maximum recording time P3 = 60 seconds, and minimum recording stop time P4 = 540 seconds, then the maximum recording time per day out of the 10 hours of operating time will be 10 hours x 60 seconds / (60 seconds + 540 seconds) = 1 hour, which means that the recording data capacity can be reduced to less than 1 / 24th of the amount required when recording continuously 24 hours a day.

[0031] The detailed processing of the recording start determining section 121 and the recording stop determining section 122 included in the recording control section 102 will be made clear in the section on operation.

[0032] The recording unit 103 records the sound recording data that the sound recording control unit 102 has determined to be recorded (the sound signals picked up by the microphone arrays MA1 and MA2, and the processed output sound).

[0033] (A-2) Operation of the embodiment Next, the operation of the voice processing device 100 of this embodiment having the above configuration will be described.

[0034] 3 is a flowchart showing the recording process of the audio processing device according to the embodiment. The series of processes in FIG. 3 starts when the current time reaches recording start time P1.

[0035] <Step S101> The recording control unit 102 determines whether or not a minimum recording stop time P4 has elapsed since the previous recording (recording stopped in step S105, described later), and proceeds to the next step S102 only if it has elapsed. Naturally, if operation starts and there is no record of a previous recording, the process proceeds to the next step S102 without making a determination. Also, as a variant, step S101 may be omitted as necessary.

[0036] <Step S102> The recording start determination unit 121 determines whether to start recording by detecting a recording start trigger, which can be, for example, (1) a user's operation of the terminal (for example, pressing a start use button via the touch panel display D), (2) detection of a user's speech in the processed output sound, or (3) detection of an operator's speech from the speaker SP.

[0037] Regarding (2) above, the processed output sound can be subjected to, for example, the voice presence / silence determination described in Patent Document 3, and the voice presence detection can be used as a trigger. Similarly, regarding (3) above, the output sound from the speaker SP can be subjected to, for example, the voice presence / silence determination described in Patent Document 3, and the voice presence detection can be used as a trigger.

[0038] As another alternative, for (2) above, it is possible to use the detection of voice as a trigger by performing the voice presence / absence determination described in Patent Document 4 on the processed output sound. Similarly, for (3) above, it is possible to use the detection of voice as a trigger by performing the voice presence / absence determination described in Patent Document 4 on the output sound from the speaker SP.

[0039] <Step S103> The recording control unit 102 starts recording the acoustic signals picked up by the microphone arrays MA1 and MA2 and the processed output sound in the recording unit 103. The recording of the audio data in the recording unit 103 continues until recording is stopped in step S105, which will be described later.

[0040] <Step S104> The recording stop determination unit 122 determines whether to stop recording by detecting a recording stop trigger, which may be, for example, (1) the user not speaking for a certain period of time, or (2) the maximum recording time P3 has elapsed since the start of recording.

[0041] The above (1) can be detected by combining hangover processing with the voice activity determination method described in Patent Document 3. An example of hangover processing in recording stop determination unit 122 will be described below with reference to Figs. 4 and 5.

[0042] FIG. 4 is a diagram showing the processing rules for the parameters (SP, TIM) involved in the hangover process of the recording stop determination unit 122. In FIG.

[0043] Here, it is assumed that the recording stop determination unit 122 has a timer variable TIM for measuring the hangover period. In the following, the frame period (frame interval) when the audio processing device 100 processes an audio signal is represented as PRD. In the following, the period set as the hangover period in the recording stop determination unit 122 is represented as HT. Here, although not particularly limited, it is assumed that the frame period PRD = 16 ms. Similarly, it is assumed that the hangover period HT = 10,000 [ms] (= 10 [sec]).

[0044] As shown in FIG. 4, the recording stop determination unit 122 sets the timer variable TIM to HT while the result of the voice activity determination process in the voice activity detection process is a voice activity section (VAD=1), and sets the voice activity section determination result to a voice activity section (SP=1).

[0045] Also, as shown in FIG. 4, if the result of the voice / silence determination process in the voice section detection process is a silent section (VAD=0) and the timer variable TIM is greater than 0 (TIM>0), the recording stop determination unit 122 continues the process of repeatedly subtracting the timer variable TIM according to the elapsed time (the process of measuring the hangover period HT using the timer variable TIM), and keeps the speech section determination result as a speech section (SP=1).

[0046] Furthermore, as shown in FIG. 4, if the result of the voice / non-voice determination process in the voice section detection process is a silent section (VAD=0) and the timer variable TIM is 0 or less (TIM≦0), the recording stop determination unit 122 determines the speech section determination result as a non-speech section (SP=0).

[0047] FIG. 5 is a timing chart showing an example of hangover processing by the recording stop determination unit 122.

[0048] When the recording stop decision unit 122 performs hangover processing according to the rules shown in FIG. 4, the parameters (VAD, TIM, SP) transition as shown in FIG.

[0049] In FIG. 5, timings T101 to T105 represent timings in time series.

[0050] [Timing T101] In FIG. 5, it is assumed that at timing T101, the result of the voice activity determination process in the voice activity detection process changes from a silent section (VAD=0) to a voice activity section (VAD=1). At this time, the recording stop determination unit 122 determines that a speech section has been detected, and changes the speech section determination result from a non-speech section (SP=0) to a speech section (SP=1). Also, at this time, the recording stop determination unit 122 sets the timer variable to HT because the speech section determination result is now a speech section (SP=1). At this time, the speech section determination result becomes a speech section (SP=1) (in other words, it is recording start timing).

[0051] [Timing T102] Assume that after timing T101, the result of the voice activity determination process remains in the voice activity section (VAD=1) state (i.e., the speech section), and at timing T102, the result of the voice activity determination process switches from the voice activity section (VAD=1) to the voice activity section (VAD=0). The recording stop determination unit 122 then starts decrementing the timer variable TIM (measuring the hangover period HT). At this time, the recording stop determination unit 122 repeats the process of subtracting (decrementing) the value of the frame period PRD from the timer variable TIM (TIM=TIM-PRD) for each frame period PRD.

[0052] [Timing T103] Here, it is assumed that a new user voice is detected before the maximum period (10 seconds) of the hangover period HT has elapsed from timing T102 (the result of the voice activity determination process is assumed to have switched from a silent period (VAD=0) to a voice activity period (VAD=1) again). In this case, the subtracted timer variable TIM is set to HT again.

[0053] [Timing T104] After timing T103, the result of the voice activity determination process remains in the voice activity section (VAD=1) state (i.e., the speech section), and at timing T104, the result of the voice activity determination process switches from the voice activity section (VAD=1) to the silent section (VAD=0). Then, the recording stop determination unit 122 starts subtracting the timer variable TIM (measuring the hangover period HT).

[0054] [Timing T105] When 10 seconds have passed since timing T104, the timer variable TIM becomes 0 (TIM≦0). In Fig. 4, the timing when the timer variable TIM becomes 0 (the timing when the hangover period HT has passed since timing T104) is designated as timing T105. At timing T105, the recording stop determination unit 122 switches the speech section determination result from a speech section (SP=1) to a non-speech section (SP=0).

[0055] At this time, the speech section determination result is a non-speech section (SP=0), so recording stop determination section 122 determines to "stop recording."

[0056] <Step S105> When it is determined in step S104 above that recording should be stopped, the recording control unit 102 stops recording the acoustic signals picked up by each microphone array MA1, MA2 and the processed output sound, which began in step S103 above, into the recording unit 103.

[0057] <Step S106> The recording control unit 102 determines whether the current time has passed the recording stop time P2, and if not, returns to step S101 described above, while if it has passed, ends the series of processes.

[0058] (A-3) Effects of the embodiment According to this embodiment, the following effects are achieved.

[0059] Because the audio processing device 100 records both the microphone array input signal (sound input from each microphone) and the processed output sound, it can make the following judgments (1) and (2) in later failure analysis. (1) If the sound input from each microphone no longer changes from a certain value, it can be determined that there is a problem with the microphone hardware, including a broken wire or a loose connector. (2) Furthermore, if the sound input from each microphone of the microphone array can be heard normally, but sound breaks or other issues occur in the processed output sound, it can be assumed that there is a problem with the audio processing parameters or algorithm. After adjusting the parameters and improving the algorithm, it is also possible to check whether the processed output sound can be improved by performing a simulation using sound recorded from the input microphones.

[0060] Furthermore, as shown in the processing of steps S101 to S106 above, efforts are made to compress the audio data stored in the audio processing device 100, so the capacity of the recording unit 103 is not strained compared to when the data is simply recorded.

[0061] (B) Other embodiments The present invention is not limited to the above-described embodiment, and may include modified embodiments such as those exemplified below.

[0062] (B-1) In the above embodiment, the acoustic signals picked up by the microphone arrays MA1 and MA2 and the processed output sound are recorded, but other sounds may also be recorded. For example, the output sound from the speaker SP or the sound in the processing process may be recorded.

[0063] (B-2) One or more of the acoustic signals picked up by the microphone arrays MA1 and MA2, the processed output sound, the speaker output sound, and the sound in the processing process may be selected and recorded.

[0064] (B-3) In the above embodiment, no particular mention was made of the recording format of the audio data stored in the recording unit 103, but the audio data stored in the recording unit 103 may be encrypted to reduce the risk of personal data being leaked to the outside.

[0065] (B-4) In the above embodiment, the voice processing device 100 is applied to an information terminal such as a ticket vending machine, but the present invention may also be applied to microphone products, speaker microphone products, and the like.

[0066] (B-5) In the above embodiment, an example was shown in which recorded data was compressed using parameters such as recording start time P1, recording stop time P2, maximum recording time length P3, and minimum recording stop time length P4. However, some of these parameters may be used (in other words, some of these parameters may be omitted). For example, the minimum recording stop time length P4 may be omitted (step S101 above may be omitted). However, in this case, if many users alternately use the voice processing device 100, the amount of recorded data will be enormous. Therefore, appropriate adjustments may be made, such as shortening the hangover period HT in the above-mentioned hangover process to less than 10 seconds, or not performing the hangover process at all. In any case, there are no particular limitations on which of the parameters to use, what values ​​to set, or whether to apply the hangover process in the recording stop determination. [Explanation of symbols]

[0067] 100...audio processing device, 101...audio processing unit, 102...recording control unit, 103...recording unit, 121...recording start determination unit, 122...recording stop determination unit, D...touch panel display, M...microphone, MA (MA1, MA2)...microphone array, P1...recording start time, P2...recording stop time, P3...maximum recording time, P4...minimum recording stop time, SP...speaker.

Claims

1. an audio processing means for performing predetermined audio processing on an input signal input from a microphone and outputting a processed output sound; a recording means for recording the input signal and the processed output sound; a recording control means for determining whether or not to record the input signal and the processed output sound on the recording means based on a predetermined condition; 10. A voice processing device comprising:

2. The recording control means a recording start determination means for determining whether or not a recording start trigger, which is a timing for starting recording of the input signal and the processed output sound in the recording means, is detected; a recording stop determination means for determining whether to stop recording based on whether a recording stop trigger, which is timing to stop recording, is detected after starting recording of the input signal and the processed output sound in the recording means after detecting the recording start trigger; 2. The audio processing device according to claim 1, further comprising:

3. 3. The audio processing device according to claim 2, wherein the recording start trigger is when a predetermined input is received from a user via an input means, when the user's voice is input to the microphone, or when voice is output from a speaker.

4. 3. The audio processing device according to claim 2, wherein the recording stop trigger is when the maximum recording time, which is the time limit for recording the input signal and the processed output sound in one recording, is exceeded, or when there is no audio input to the microphone for a certain period of time.

5. The recording stop determination means a sound segment detection means for detecting a sound segment containing speech in the processed output sound; a speech interval detection means for detecting an utterance interval in which at least the processed output sound is a sound interval; a voice input determination means for performing a hangover process to extend the state of the speech section for a predetermined hangover period after the processed output sound has changed from a voice section to a silent section, and determining that there has been no voice input to the microphone for a certain period of time if the silent section continues even after the state of the speech section has been extended to the maximum extent; 5. The audio processing device according to claim 4, further comprising:

6. 6. The audio processing device according to claim 5, wherein the recording control means limits the recording means from recording the next input signal and the processed output sound until a predetermined time has elapsed since the input signal and the processed output sound were recorded once in the recording means.

7. 7. The audio processing device according to claim 6, wherein the recording control means determines whether or not to record the input signal and the processed output sound in the recording means only during a period from a preset recording start time to a preset recording stop time.

8. Computer, an audio processing means for performing predetermined audio processing on an input signal input from a microphone and outputting a processed output sound; a recording means for recording the input signal and the processed output sound; a recording control means for determining whether or not to record the input signal and the processed output sound on the recording means based on a predetermined condition; A speech processing program characterized by functioning as follows.

9. 1. A voice processing method for use in a voice processing device, comprising: The audio processing device has an audio processing means, a recording means, and a recording control means, the audio processing means performs predetermined audio processing on an input signal input from a microphone and outputs a processed output sound; the recording means records the input signal and the processed output sound; The recording control means determines whether or not to record the input signal and the processed output sound on the recording means based on a predetermined condition.

1. A sound processing method comprising:

Citation Information

Patent Citations

  • Sound / soundless determination device, sound / soundless determination method, and sound / soundless determination program

    JP2011070084A

  • Sound collection device and program

    JP2017183902A

  • Voice processing device

    JP2021033030A

  • Voice detection device, voice detection program, and voice detection method

    JP2022032721A

  • Sound collection device, sound collection program, sound collection method, and sound collection system

    JP7207170B2