Image recording device and image recording method
The image recording device dynamically sets the focus area based on sound direction to enhance image quality at accident scenes or road rage incidents, addressing the limitations of existing devices by optimizing encoding and reducing data size.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- JVC KENWOOD CORP
- Filing Date
- 2021-12-24
- Publication Date
- 2026-05-26
AI Technical Summary
Existing vehicle-mounted image recording devices lack the ability to dynamically set the area of focus based on the direction of arrival of sounds, such as impact sounds, horns, sirens, or screams, to enhance image quality at accident scenes or road rage incidents, as they rely on predetermined sound triggers rather than dynamic sound direction detection.
An image recording device equipped with a sound acquisition unit, sound direction detection unit, focus area setting unit, and image encoding unit to detect sound direction and dynamically set the region of interest for higher image quality in the direction of sound arrival, performing different encoding processes on the focus area and other areas.
This approach allows for clearer recording of accident scenes or road rage incidents by enhancing image quality in the direction of sound arrival, reducing data size by optimizing encoding, and ensuring high-quality capture of relevant areas while minimizing unnecessary data storage.
Smart Images

Figure 0007865007000001 
Figure 0007865007000002 
Figure 0007865007000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image recording device and an image recording method. [Background technology]
[0002] In recent years, progress has been made in developing image recording devices to accurately record the circumstances surrounding accidents and other incidents.
[0003] Patent Document 1 discloses an image processing device that sets a region of interest (ROI) for an image frame of a moving image captured by an imaging means, detects the region of a moving object in this image frame, determines whether at least a part of the detected region of the moving object is included in the region of interest, and controls the encoding of the moving image inside and outside the region of interest according to the result of this determination. [Prior art documents] [Patent Documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2019-134323 [Overview of the Initiative] [Problems that the invention aims to solve]
[0005] Vehicle-mounted image recording devices, such as drive recorders, continuously capture images of the vehicle's surroundings and record images and video (hereinafter simply referred to as "images"). In recent years, the image quality of these image recording devices has increased, and consequently, the size of the image files has also increased.
[0006] In the technology disclosed in Patent Document 1, when a moving object is detected in the area of interest, the detection of a predetermined sound (for example, the sound of a window breaking, or a person shouting) triggers encoding to make the area of interest high-resolution. In this way, by selectively making the area of interest high-resolution and relatively lowering the resolution of areas other than the area of interest, it is possible to improve the resolution of the necessary areas while reducing the size of the image file.
[0007] In response to this, the inventors considered detecting the direction of arrival of sounds (e.g., impact sounds, horns, sirens / brake sounds, screams / shouts) around a vehicle equipped with an image recording device such as a drive recorder, and dynamically setting the area corresponding to the detected direction of arrival of the sound as the area of focus. This allows for higher image quality in the area where the sound (e.g., impact sounds, horns, sirens / brake sounds, screams / shouts) around the vehicle is coming from (the location where an abnormality such as an accident scene or a road rage incident occurred), compared to areas outside the area of focus, thereby enabling clearer recording of the accident scene or the location where an abnormality such as a road rage incident occurred. Note that the sounds around the vehicle are not limited to impact sounds, horns, sirens / brake sounds, screams / shouts, but may also include other sounds (such as abnormal sounds).
[0008] However, Patent Document 1 merely describes a predetermined sound as a trigger for encoding a pre-set area of focus (which is not dynamically set) to achieve high image quality, and does not disclose anything about detecting the direction of arrival of sounds around the vehicle and dynamically setting the area corresponding to the detected direction of arrival of the sound as the area of focus.
[0009] The present invention has been made in view of the above points, and aims to provide an image recording device and an image recording method that can detect the direction of arrival of sound around a vehicle and dynamically set the region corresponding to the detected direction of arrival of the sound as a region of focus. [Means for solving the problem]
[0010] The image recording device according to the present invention comprises: an image acquisition unit that acquires an image of the area around a vehicle; a sound acquisition unit that acquires sound from the area around the vehicle; a sound direction detection unit that detects the direction from which the sound is coming; a focus area setting unit that sets a region of the image corresponding to the direction from which the sound is coming as a focus area; and an image encoding unit that performs a first encoding process on the focus area of the image and a second encoding process on the areas of the image other than the focus area.
[0011] The image recording method according to the present invention comprises: an image acquisition step of acquiring an image of the area around a vehicle; a sound acquisition step of acquiring sound from the area around the vehicle; a sound direction detection step of detecting the direction from which the sound is coming; a focus area setting step of setting a region of the image corresponding to the direction from which the sound is coming as a focus area; and an image encoding step of performing a first encoding process on the focus area of the image and a second encoding process on the areas of the image other than the focus area. [Effects of the Invention]
[0012] According to this disclosure, it is possible to provide an image recording device and an image recording method that can detect the direction of arrival of sound around a vehicle and dynamically set the region corresponding to the detected direction of arrival of the sound as a region of focus. [Brief explanation of the drawing]
[0013] [Figure 1] This is a block diagram showing an example configuration of the image recording device 1 according to Embodiment 1. [Figure 2] This is an example of the setup of camera 19 and microphone 20. [Figure 3] This is a flowchart of the operation (image recording method) of the image recording device 1 according to Embodiment 1. [Figure 4] This is an example of an image I of the surroundings of the vehicle V. [Figure 5] This is a flowchart illustrating a specific example of encoding process 2. [Figure 6] This is an example of the volume of sound (detected volume) acquired by the sound acquisition unit 12. [Figure 7] It is a flowchart of a specific example of the attention area setting process. [Figure 8] It is a flowchart of a specific example of the encoding process 1. [Figure 9] It is a flowchart of other operations of the image recording apparatus 1 (a modified example of the image recording method). [Figure 10] It is a block diagram showing a configuration example of the image recording apparatus 1A according to the second embodiment. [Figure 11] It is another installation example of the camera 19 and the microphone 20. [Figure 12] It is a flowchart of the operation (image recording method) of the image recording apparatus 1A according to the second embodiment. [Figure 13] It is a flowchart of a specific example of the attention image setting process. [Figure 14] It is a hardware configuration example of the control device of the image recording apparatus according to the first and second embodiments.
Embodiments for Carrying Out the Invention
[0014] Hereinafter, the present invention will be described through embodiments of the invention, but the invention according to the claims is not limited to the following embodiments. Also, not all of the configurations described in the embodiments are essential as means for solving the problems. For the sake of clarity of explanation, the following description and drawings are appropriately omitted and simplified. In each drawing, the same reference numerals are assigned to the same elements, and redundant explanations are omitted. <First Embodiment> Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0015] FIG. 1 is a block diagram showing a configuration example of the image recording apparatus 1 according to the first embodiment. As shown in FIG. 1, the image recording apparatus 1 according to the present embodiment includes an image acquisition unit 11, a sound acquisition unit 12, a sound direction detection unit 13, an attention area setting unit 14, an image encoding unit 15, a recording control unit 16, a sound recognition unit 17, and a recording unit 18.
[0016] The image recording device 1 is a device mounted on a vehicle V (see Figure 2; hereinafter also referred to as the vehicle V), and is, for example, a drive recorder. The image recording device 1 records images acquired by the image acquisition unit 11 in the recording unit 18. The images recorded at this time are encoded in the image encoding unit 15. In this embodiment, the recording unit 18 is, for example, a recording medium such as an SD card, SSD (Solid State Drive), HDD (hard disk drive), or memory. The recording unit 18 may be built into the image recording device 1, or it may be configured to be removable from the image recording device 1. Alternatively, the recording unit 18 may be provided outside the image recording device 1. Note that the encoded images may not be recorded in the recording unit 18, but may be transmitted externally (for example, to a management company such as a taxi company) via a communication unit (not shown).
[0017] The image acquisition unit 11 acquires an image of the area surrounding the vehicle V. The area surrounding the vehicle V refers, for example, to the front, rear, sides, top, or bottom of the vehicle V. In Figure 4, symbol I is an example of an image taken of the area surrounding the vehicle V (for example, the front). Hereinafter, an image taken of the area surrounding the vehicle V (for example, the front) will be referred to as image I. Image I may include only the area surrounding the vehicle V, or it may include the area surrounding the vehicle V and the interior of the vehicle V. The image acquisition unit 11 may include a camera, or it may be an input circuit that inputs or receives signals (image signals) transmitted from the camera.
[0018] Figure 2 shows an example of the installation of camera 19 and microphone 20.
[0019] As shown in Figure 2, the camera 19 is installed, for example, inside the vehicle V (for example, in the center in the width direction of the vehicle) and captures images of the area in front of the vehicle V through the windshield. The optical axis AX of the camera 19 extends in the longitudinal direction of the vehicle. The camera 19 may also be installed outside the vehicle V.
[0020] The sound acquisition unit 12 (sound detection unit) acquires (detects) sounds from the surroundings of the vehicle V. The sound acquisition unit 12 may include a microphone, or it may be an input circuit that inputs or receives signals (sound signals) transmitted from the microphone.
[0021] As shown in Figure 2, it is desirable to install the microphones 20 at each of the following locations on the vehicle V (for example, a total of 4 locations): above, below, to the left, and to the right. In this case, it is desirable to install two microphones 20 symmetrically with respect to the optical axis AX of the camera 19 in the left-right direction, and the remaining two microphones 20 symmetrically with respect to the optical axis AX of the camera 19 in the up-down direction.
[0022] The sound acquisition unit 12 acquires sound (sound signals) transmitted from each microphone 20 (microphone array). The sound acquisition unit 12 transmits the volume of the sound acquired from each microphone 20 and information on the installation position of each microphone 20 to the sound direction detection unit 13. The sound acquisition unit 12 also transmits the sound (sound signals) transmitted from each microphone 20 (microphone array) to the sound recognition unit 17.
[0023] The sound direction detection unit 13 detects the direction of arrival of the sound acquired by the sound acquisition unit 12 based on the data transmitted from the sound acquisition unit 12 (the delay amount and sound pressure of multiple sounds transmitted from each microphone 20) and information on the installation position of each microphone 20. The direction of arrival of the sound corresponds to the optical axis AX of the camera 19, and is located on P in the image I (acquired by the image acquisition unit 11). AX (The origin on the XY coordinate system; see Figure 4) is used as the reference point for detection. For example, from the sound of two microphones 20 placed to the left and right of the optical axis AX of the camera 19, P on image I is detected. AX The system detects that the direction of sound arrival is to the right, and similarly, from the sound of two microphones 20 placed vertically relative to the optical axis AX of the camera 19, the P on image I is determined. AX The upward direction is detected as the direction of sound arrival. The direction of sound arrival may be detected in more detail based on the delay amount and sound pressure difference. In Figure 4, the symbol P. ROIThe black circle indicated by (coordinate position on the XY coordinate system) represents the direction of sound arrival (sound source direction) relative to image I acquired by the image acquisition unit 11. Specific examples of sound direction detection will be described later.
[0024] For detecting the direction of sound arrival, known techniques for identifying or estimating the direction of a sound source, such as beamforming technology using a microphone array, may be used. In this case, it is preferable to perform noise reduction processing (removal or reduction of road noise and wind noise) on the sound acquired by the sound acquisition unit 12 as a preliminary step.
[0025] The Region of Interest Setting Unit 14 executes the Region of Interest Setting Process. The Region of Interest Setting Process is the process of setting a region of interest (ROI) in the image I acquired by the image acquisition unit 11 that corresponds to the direction of arrival of the sound, based on the direction of arrival of the sound detected by the sound direction detection unit 13. This makes it possible to dynamically set a region of interest as an ROI that corresponds to the direction of arrival of sounds around the vehicle (for example, impact sounds, horns, sirens / brake sounds, screams / shouts) (locations where abnormalities such as accident sites or road rage incidents have occurred). Specific examples of setting the Region of Interest ROI will be described later. The Region of Interest ROI refers to a region in the image I acquired by the image acquisition unit 11 that corresponds to the direction of arrival of the sound detected by the sound direction detection unit 13. In addition, the Region of Interest Setting Unit 14 may set the Region of Interest ROI based on the direction of arrival of the sound detected by the sound direction detection unit 13 when the sound recognition unit 17, described later, recognizes that the sound acquired by the sound acquisition unit 12 is a pre-registered sound.
[0026] For example, the ROI (Region of Interest) may be one (or more) small regions corresponding to the direction of sound arrival, from among multiple small regions obtained by dividing the image I acquired by the image acquisition unit 11, for example, multiple rectangular small regions divided into 9 sections (3x3 divisions) or 16 sections (4x4 divisions). However, the ROI may be a region of a shape other than a rectangle (for example, circular, elliptical). Furthermore, the ROI may be the entire image I acquired by the image acquisition unit 11.
[0027] The size of the ROI (Region of Interest) may be changed depending on the volume of the sound acquired by the sound acquisition unit 12. For example, if the volume of the sound acquired by the sound acquisition unit 12 is high, there is a possibility that the sound source is closer to the vehicle compared to when the volume is low. Therefore, if the volume of the sound acquired by the sound acquisition unit 12 (detected volume) is greater than a pre-registered threshold, the size of the ROI may be increased. Conversely, if the volume of the sound acquired by the sound acquisition unit 12 (detected volume) is lower than a pre-registered threshold, the size of the ROI may be decreased.
[0028] The image encoding unit 15 performs image encoding processing (for example, encoding processing 1 or encoding processing 2). This encoding may be any encoding format such as MPEG, JPEG, or H.265.
[0029] Encoding process 1 refers to the encoding process that is executed when a region of interest (ROI) is set. In encoding process 1, for the region of interest (ROI) among the image I acquired by the image acquisition unit 11, a process is performed (hereinafter referred to as the first encoding process) that assigns more codes to the region of interest than when the second encoding process described later is performed, for example, the compression ratio low Encoding (first encoding) is performed by executing a process to control quantization parameters, such as setting the DCT coefficient to increase the amount of code. For areas of the image I acquired by the image acquisition unit 11 other than the region of interest (ROI), a process is performed to assign fewer codes than when the first encoding process is performed (hereinafter referred to as the second encoding process), for example, by adjusting the compression ratio. highEncoding (second encoding) is performed by executing a process to control quantization parameters, such as setting the DCT coefficient to reduce the amount of code. Note that the bitrate when the first encoding process is performed will be higher than the bitrate when the second encoding process is performed. As a result, the region of interest (ROI) will have higher image quality than the region of interest (ROI), so the direction from which sounds around the vehicle (e.g., impact sounds, horns, sirens / brake sounds, screams / shouts) are coming (locations where abnormalities such as accident scenes or road rage incidents occur) can be recorded in high quality. Furthermore, by reducing the amount of code allocated to the region of interest (ROI), the amount of data required for recording can be reduced. The degree of quantization parameter control in the region of interest (ROI) may be changed depending on the volume of sound acquired by the sound acquisition unit 12 (detected volume). For example, if the volume of sound acquired by the sound acquisition unit 12 is high, it may indicate a major accident or an accident that occurred very close to vehicle V. For this reason, if the volume of sound acquired by the sound acquisition unit 12 (detected volume) is greater than a pre-registered threshold, the quantization parameters may be controlled to have a larger amount of code. Conversely, if the volume of the sound acquired by the sound acquisition unit 12 (detected volume) is smaller than a pre-registered threshold, the quantization parameters may be controlled to result in a smaller coding amount.
[0030] Encoding process 2 refers to the encoding process that is executed when no region of interest (ROI) is set. In encoding process 2, the second encoding process is performed on the image I (entire image) acquired by the image acquisition unit 11.
[0031] The recording control unit 16 records the image encoded by the image encoding unit 15 to the recording unit 18. Since the first encoding process is performed on the region of interest (ROI), the direction from which sounds around the vehicle (e.g., impact sounds, horns, sirens / brake sounds, screams / shouts) are coming (locations where abnormalities such as accident scenes or road rage incidents occur) can be recorded in high quality.
[0032] Images outside the region of interest (ROI) or those without a defined ROI undergo a second encoding process and therefore will not be of high quality. By not unnecessarily enhancing the quality of areas outside the ROI, the amount of data required for recording can be reduced.
[0033] If the recording control unit 16 has not acquired (detected) sound after the sound acquisition unit 12 has acquired sound (detected sound) but a predetermined period of time has elapsed, it will restore the quantization parameters of the region of interest (ROI) to their original values and perform a second encoding process. The state of not acquiring sound may be a state in which sounds with a volume greater than a pre-registered threshold are not acquired (or not detected).
[0034] The sound recognition unit 17 performs sound recognition processing on the sound acquired by the sound acquisition unit 12 and performs recognition processing to determine if the sound is a sound that has been registered in advance. Pre-registered sounds include, for example, abnormal sounds such as collision sounds, horns, sirens, brake sounds, screams, and shouts, and are provided in the sound recognition unit 17 as a recognition dictionary. The sound recognition unit 17 is not necessarily a required component in Embodiment 1.
[0035] Next, the operation (image recording method) of the image recording device 1 according to Embodiment 1 will be explained using the flowchart shown in Figure 3.
[0036] Figure 3 is a flowchart of the operation (image recording method) of the image recording device 1 according to Embodiment 1.
[0037] First, the image acquisition unit 11 acquires an image of the area around the vehicle V (step S10). Here, as shown in Figure 4, it is assumed that an image I, taken from the front of the vehicle V, is acquired as the image of the area around the vehicle V. Figure 4 is an example of image I of the area around the vehicle V.
[0038] Next, it is determined whether or not a region of interest (ROI) has already been set (step S11).
[0039] If the result of the determination in step S11 is determined to be that the region of interest (ROI) has not been set (step S11: NO), the processes from step S12 onwards will be executed.
[0040] First, the sound acquisition unit 12 acquires (detects) sounds from around the vehicle V (step S12). Specifically, the sound acquisition unit 12 acquires sounds (sound signals) transmitted from each of the microphones 20 (a total of 4) as sounds from around the vehicle V.
[0041] Next, the sound recognition unit 17 performs sound recognition processing on the sound acquired by the sound acquisition unit 12 in step S12 (step S13), and determines whether the sound is a pre-registered sound (an abnormal sound such as a collision sound, horn, siren, brake sound, scream, shout, etc.) (step S14). Note that steps S13 and S14 are not mandatory processes, and the system may proceed from step S12 to S15.
[0042] If, as a result of the determination in step S14, it is determined that the sound acquired by the sound acquisition unit 12 in step S12 is not a pre-registered sound (abnormal sound) (step S14: NO), the image encoding unit 15 executes encoding process 2 (step S21).
[0043] Here, we will explain a specific example of encoding process 2.
[0044] Figure 5 is a flowchart of a specific example of encoding process 2.
[0045] As shown in Figure 5, the recording control unit 16 performs a second encoding process on the image I (entire image) acquired by the image acquisition unit 11 in step S10 (step S211).
[0046] Let's continue by returning to Figure 3.
[0047] Once encoding process 2 (step S21) is executed, the recording control unit 16 then records the image encoded in step S21 into the recording unit 18 (step S22). Specifically, the recording control unit 16 records the image that has undergone the second encoding process on the image I (entire image) acquired by the image acquisition unit 11 in step S10 into the recording unit 18. The encoded image is recorded as a file.
[0048] Next, the process returns to step S10, and the subsequent steps are repeatedly executed.
[0049] On the other hand, if the result of the determination in step S14 is determined to be a sound that has been registered in advance (collision sound, horn, siren, brake sound, scream, shout, or other abnormal sound) (step S14: YES), then the processes from step S15 onwards are executed.
[0050] First, the sound direction detection unit 13 detects the direction of the sound acquired by the sound acquisition unit 12 in step S12 (step S15) based on the data transmitted from the sound acquisition unit 12 (information on the volume of the sound acquired by the sound acquisition unit 12 in step S12 and the installation position of each microphone 20).
[0051] Here, we will explain a specific example of detecting the direction of sound arrival.
[0052] For example, as shown in Figure 6, suppose that the volume of sound (detected volume) acquired by the sound acquisition unit 12 in step S12 is obtained. Figure 6 is an example of the volume of sound (detected volume) acquired by the sound acquisition unit 12. In Figure 6, the X-axis + direction refers to the microphone 20 installed to the right of the optical axis AX of the camera 19. Similarly, in Figure 6, the X-axis - direction refers to the microphone 20 installed to the left of the optical axis AX of the camera 19, the Y-axis + direction refers to the microphone 20 installed upwards of the optical axis AX of the camera 19, and the Y- direction refers to the microphone 20 installed downwards of the optical axis AX of the camera 19. 3, 1, 5, and 2 in Figure 6 represent the volume of sound acquired from each microphone 20.
[0053] In this case, as shown in FIG. 4, the sound direction detection unit 13 uses the arrival direction of the sound acquired by the sound acquisition unit 12 in step S12 (the position of the sound detection direction in the image I), i.e., the center of the image I, that is, P XAX as a reference to detect the position at +2 in the X-axis direction (3 - 1 = +2) and the position at +3 in the Y-axis direction (5 - 2 = +3). This detected position is represented by the black circle indicated by the symbol P ROI in FIG. 4. Hereinafter, this detected position will be referred to as the detection position P ROI In FIG. 4, the X-axis passes through the center of the image I, that is, P AX and extends in the left-right direction. On the other hand, the Y-axis passes through the center of the image I, that is, P AX and extends in the up-down direction. Also, according to the arrival time of the sound acquired from each microphone 20, the delay amount of the arrival time of the sound between the microphones may be calculated, and the arrival direction of the sound acquired by the sound acquisition unit 12 may be detected using the speed of sound and the distance between the microphones.
[0054] Next, the attention area setting unit 14 executes an attention area setting process (step S16).
[0055] Here, a specific example of the attention area setting process (a specific example of setting the attention area ROI) will be described.
[0056] FIG. 7 is a flowchart of a specific example of the attention area setting process.
[0057] As shown in FIG. 7, first, the attention area setting unit 14 searches for the area corresponding to the arrival direction of the sound (the arrival direction of the sound detected in step S15) among the plurality of rectangular small areas (refer to the nine rectangular small areas partitioned by dotted lines in FIG. 4) obtained by dividing the image I acquired by the image acquisition unit 11 in step S10 (step S161). ROI Here, as the area corresponding to the arrival direction of the sound, it is assumed that the small area including the detection position P
[0058] (the upper right small area in FIG. 4) among the plurality of rectangular small areas (refer to the nine rectangular small areas partitioned by dotted lines in FIG. 4) is searched. ROI is searched.
[0059] Next, the area of interest setting unit 14 sets the area found in step S161 as the area of interest ROI (step S162).
[0060] Here, the detection position P found in step S161 is located here. ROI The small area in the upper right corner, including the specified region, is assumed to be the region of interest (ROI).
[0061] Thus, the region of interest setting process allows for the dynamic setting of a region of interest (ROI) that corresponds to the direction from which sounds around the vehicle (e.g., impact sounds, horns, sirens / brake sounds, screams / shouts) are coming (locations where abnormalities such as accident sites or road rage incidents have occurred).
[0062] The size of the ROI of interest may be changed depending on the volume of the sound acquired by the sound acquisition unit 12 in step S12. For example, if the volume of the sound acquired by the sound acquisition unit 12 is high, there is a possibility that the sound source is closer to the vehicle compared to when the volume is low. Therefore, if the volume of the sound acquired by the sound acquisition unit 12 in step S12 (detected volume) is greater than a pre-registered threshold, the size of the ROI of interest may be increased. Conversely, if the volume of the sound acquired by the sound acquisition unit 12 in step S12 (detected volume) is less than a pre-registered threshold, the size of the ROI of interest may be decreased.
[0063] Let's continue by returning to Figure 3.
[0064] Once the region of interest (ROI) is set (step S16), the image encoding unit 15 then executes the encoding process 1 (step S17).
[0065] Here, we will explain a specific example of encoding process 1.
[0066] Figure 8 is a flowchart of a specific example of encoding process 1.
[0067] As shown in Figure 8, first, the image encoding unit 15 performs a first encoding process on the region of interest (ROI) of the image I acquired by the image acquisition unit 11 in step S10 (step S171). Meanwhile, the image encoding unit 15 performs a second encoding process on the region of image I other than the region of interest (ROI) of the image I acquired by the image acquisition unit 11 in step S10 (step S172).
[0068] The degree of quantization parameter control may be changed depending on the volume of the sound acquired by the sound acquisition unit 12 in step S12 (detected volume). For example, if the volume of the sound acquired by the sound acquisition unit 12 is high, there is a possibility of a major accident. Therefore, if the volume of the sound acquired by the sound acquisition unit 12 in step S12 (detected volume) is greater than a pre-registered threshold, the quantization parameters may be controlled to increase the amount of code allocated. Conversely, if the volume of the sound acquired by the sound acquisition unit 12 in step S12 (detected volume) is lower than a pre-registered threshold, the quantization parameters may be controlled to decrease the amount of code allocated.
[0069] Let's continue by returning to Figure 3.
[0070] Once encoding process 1 (step S17) is performed, the recording control unit 16 then records the image encoded by the image encoding unit 15 in the recording unit 18 (step S18). Specifically, the recording control unit 16 performs a first encoding process on the region of interest (ROI) of image I and records it in the recording unit 18. On the other hand, the recording control unit 16 performs a second encoding process on the region of image I other than the region of interest and records it in the recording unit 18.
[0071] In this way, the ROI of interest is recorded with higher image quality than areas outside the ROI of interest, making it possible to record the direction of sound from around the vehicle (e.g., impact sounds, horns, sirens / brake sounds, screams / shouts) (locations where abnormalities such as accident scenes or road rage incidents occur) with high image quality.
[0072] Next, it is determined whether there is silence for a predetermined period of time (step S19). Silence means that the volume is below a certain level (for example, below a predetermined threshold).
[0073] If the result of the determination in step S19 is that there is no silence for the predetermined time (step S19: NO), the process returns to step S10 and the subsequent processing is repeated.
[0074] Next, it is determined whether or not a region of interest (ROI) has already been set (step S11).
[0075] Here, since the ROI of interest has already been set, it is determined that the ROI of interest has been set (step S11: YES), and the processing from step S17 onwards is executed.
[0076] In other words, steps S10, S11:YES, S17, and S18 are repeatedly executed until it is determined in step S19 that there has been silence for a predetermined period of time.
[0077] Then, if it is determined in step S19 that there has been silence for a predetermined period of time (step S19: YES), that is, if a predetermined period of time has elapsed since the sound acquisition unit 12 acquired sound (detected sound) but no sound has been acquired (no sound has been detected), the recording control unit 16 resets the quantization parameters of the region of interest (ROI) to their original values (step S20). In other words, the setting of the region of interest (ROI) is released.
[0078] Next, the process returns to step S10, and the subsequent steps are repeatedly executed.
[0079] Furthermore, when playing back an image recorded in the recording unit 18, an indication that the ROI of interest can be identified (for example, an indication that the ROI of interest is surrounded by a specific color) may be displayed in the played-back image.
[0080] As described above, according to Embodiment 1, since it is equipped with a sound direction detection unit 13, it is possible to detect the direction of arrival of sound around the vehicle, and since it is equipped with a focus area setting unit 14, it is possible to dynamically set the area corresponding to the detected direction of arrival of sound as a focus area ROI. Furthermore, since it is equipped with an image encoding unit 15, it is possible to perform a first encoding process on the focus area ROI in image I and a second encoding process on the areas of image I other than the focus area ROI.
[0081] This allows for clearer recording of areas where abnormalities occur, such as accident scenes or road rage incidents, by highlighting the direction from which sounds around the vehicle (e.g., impact sounds, horns, sirens / brake sounds, screams / shouts) are coming from, using the ROI (Region of Interest) as the source of the ROI.
[0082] Next, a modified example will be described. In the modified example, unlike Embodiment 1 in which the region of interest and the region outside the region of interest are set separately within a single image depending on the direction of sound arrival, the entire image is set as the region of interest when the direction of sound arrival is within the field of view of the image.
[0083] Figure 9 is a flowchart of another operation of the image recording device 1 (a modified image recording method). Figure 9 corresponds to replacing steps S16 to S18 in Figure 3 with steps S16A to S18A.
[0084] The following explanation will focus on steps S16A to S18A, which are the differences from Figure 3.
[0085] When the sound direction detection unit 13 detects the direction of arrival of the sound acquired by the sound acquisition unit 12 in step S12 (step S15), it is then determined whether or not the detected direction of arrival of the sound is within the field of view of the camera 19 (step S16A).
[0086] If the determination in step S16A does not determine that the direction of the sound is within the field of view of the camera 19 (step S16A: NO), that is, if the direction of the sound is outside the field of view of the camera 19, then the processes in steps S21 and S22 are executed. Specifically, the image encoding unit 15 executes encoding process 2 (step S21), and the recording control unit 16 records the image encoded in step S21 to the recording unit 18 (step S22).
[0087] On the other hand, if the determination in step S16A determines that the direction of the sound is within the field of view of the camera 19 (step S16A: YES), the image encoding unit 15 executes encoding process 3 (step S17A). In encoding process 3, the first encoding process is performed on the entire image I acquired by the image acquisition unit 11 in step S10.
[0088] Next, the recording control unit 16 records the image encoded in step S17A to the recording unit 18 (step S18A). The encoded image is recorded as a file.
[0089] According to this modified version, if it is determined that the direction of the sound is within the field of view of camera 19 (step S16A: YES), a first encoding process is performed on the entire image I. On the other hand, if it is determined that the direction of the sound is outside the field of view of camera 19 (step S16A: NO), a second encoding process is performed on the entire image I.
[0090] This allows for high-resolution recording of images when the direction from which sounds around the vehicle (e.g., impact sounds, horns, sirens / brake sounds, screams / shouts) are coming from (locations where abnormalities such as accident scenes or road rage incidents occur) is within the field of view of camera 19, in other words, when the direction from which the sound is coming can be captured. This enables clearer recording of accident scenes, road rage incidents, and other abnormalities. <Embodiment 2> Figure 10 is a block diagram showing an example configuration of the image recording device 1A according to Embodiment 2.
[0091] Figure 10 corresponds to Figure 1 with the focus area setting unit 14 replaced by the focus image setting unit 14A.
[0092] The focus image setting unit 14A executes the focus image setting process. The focus image setting process is the process of setting the image of the camera 19 that captured the direction corresponding to the direction of sound arrival from among the multiple images I acquired by the image acquisition unit 11, which is equipped with multiple cameras 19, based on the direction of sound arrival detected by the sound direction detection unit 13. The focus image setting unit 14A may set the entire image I captured by one camera that substantially coincides with the direction of sound arrival as the focus image, or it may set the entire multiple images I captured by two or more cameras that are close to the direction of sound arrival as the focus image. The focus image is the image I acquired by the image acquisition unit 11 (for example, images captured by four cameras 19) that corresponds to the direction of sound arrival detected by the sound direction detection unit 13.
[0093] Figure 11 shows other installation examples of the camera 19 and microphone 20.
[0094] As shown in Figure 11, a total of four cameras 19 are installed to capture images of the front, rear, left side, and right side of the vehicle V. In addition, a total of four microphones 20 are installed corresponding to the cameras 19.
[0095] Next, the operation (image recording method) of the image recording device 1A according to Embodiment 2 will be explained using the flowchart shown in Figure 12.
[0096] Figure 12 is a flowchart of the operation (image recording method) of the image recording device 1A according to Embodiment 2. Figure 12 corresponds to Figure 3 with steps S11, S16-18, S21, and S22 replaced by steps S11B, S16B-18B, S21B, and S22B.
[0097] First, the image acquisition unit 11 acquires images of the area around the vehicle V (step S10). Here, it is assumed that images I (a total of 4 images) are acquired from each camera 19, capturing the area around the vehicle V (front, rear, left side, and right side).
[0098] Next, it is determined whether or not the image of interest has already been set (step S11B).
[0099] If the result of the determination in step S11B is determined to be that the image of interest has not been set (step S11B: NO), the processes from step S12 onwards will be executed.
[0100] First, the sound acquisition unit 12 acquires (detects) sounds from around the vehicle V (step S12). Specifically, the sound acquisition unit 12 acquires sounds (sound signals) transmitted from each of the microphones 20 (a total of 4) as sounds from around the vehicle V.
[0101] Next, the sound recognition unit 17 performs sound recognition processing on the sound acquired by the sound acquisition unit 12 in step S12 (step S13), and determines whether the sound is a pre-registered sound (an abnormal sound such as a collision sound, horn, siren, brake sound, scream, shout, etc.) (step S14).
[0102] If, as a result of the determination in step S14, it is determined that the sound acquired by the sound acquisition unit 12 in step S12 is not a pre-registered sound (abnormal sound) (step S14: NO), the image encoding unit 15 executes encoding process 5 (step S21B).
[0103] Encoding process 5 refers to performing a second encoding process on the total of four images I (the whole) acquired by the image acquisition unit 11 in step S10.
[0104] Next, the recording control unit 16 records the image encoded in step S21B into the recording unit 18 (step S22B). Specifically, the recording control unit 16 records the encoded image for the total of four images I (the whole) acquired by the image acquisition unit 11 in step S10 into the recording unit 18. The encoded image is recorded as a file.
[0105] Next, the process returns to step S10, and the subsequent steps are repeatedly executed.
[0106] On the other hand, if the result of the determination in step S14 is determined to be a sound that has been registered in advance (collision sound, horn, siren, brake sound, scream, shout, or other abnormal sound) (step S14: YES), then the processes from step S15 onwards are executed.
[0107] First, the sound direction detection unit 13 detects the direction of arrival of the sound acquired by the sound acquisition unit 12 in step S12 (step S15) based on the data transmitted from the sound acquisition unit 12 (the delay amount and sound pressure of multiple sounds transmitted from each microphone 20) and information on the installation position of each microphone 20.
[0108] Specifically, the sound direction detection unit 13 detects the direction from which the microphone 20 with the loudest sound volume is located as the direction from which the sound is coming. Alternatively, the sound direction detection unit 13 detects the direction from which the microphone 20 with the fastest arrival time is located as the direction from which the sound is coming. In addition, as in Embodiment 1, the direction from which the sound is coming may be detected based on the difference in sound volume and the amount of delay.
[0109] Next, the focus image setting unit 14A executes the focus image setting process (step S16B).
[0110] Here, we will explain a specific example of the process for setting the featured image.
[0111] Figure 13 is a flowchart of a specific example of the process for setting the target image.
[0112] As shown in Figure 13, first, the focus image setting unit 14A searches for the image corresponding to the direction of arrival of the sound detected by the sound direction detection unit 13 from among the multiple images I (images captured by each camera 19) acquired by the image acquisition unit 11 (step S161B).
[0113] Next, the focus image setting unit 14B sets the image found in step S161B as the focus image (step S162B). At this time, as shown in the modified example of Embodiment 1, an image in which the direction of sound arrival is included within the field of view may be set as the focus image.
[0114] Thus, the focus image setting process allows for the dynamic setting of an image as the focus image, corresponding to the direction from which sounds around the vehicle (e.g., impact sounds, horns, sirens / brake sounds, screams / shouts) are coming from (locations where abnormalities such as accident scenes or road rage incidents have occurred).
[0115] Let's continue by returning to Figure 12.
[0116] Once the image of interest is set (step S16B), the image encoding unit 15 then performs encoding process 4 (step S17B). Specifically, the image encoding unit 15 performs a first encoding process on the entire image of interest from among the four images captured by each camera 19, while performing a second encoding process on the images other than the image of interest from among the four images captured by each camera 19.
[0117] Next, the recording control unit 16 records the image encoded by the image encoding unit 15 to the recording unit 18 (step S18B). At this time, the encoded image is recorded as a file.
[0118] In this way, the image of interest is enhanced in higher resolution than other images, making it possible to record in high resolution the direction from which sounds around the vehicle (for example, impact sounds, horns, sirens / brake sounds, screams / shouts) are coming from (locations where abnormalities such as accident scenes or road rage incidents have occurred).
[0119] Next, it is determined whether or not there is silence for a predetermined period of time (step S19).
[0120] If the result of the determination in step S19 is that there is no silence for the predetermined time (step S19: NO), the process returns to step S10 and the subsequent processing is repeated.
[0121] Next, it is determined whether or not the image of interest has already been set (step S11B).
[0122] Here, since the target image has already been set, it is determined that the target image has been set (step S11B: YES), and the processing from step S17B onward is executed.
[0123] In other words, steps S10, S11B:YES, S17B, and S18B are repeatedly executed until it is determined in step S19 that there has been silence for a predetermined period of time.
[0124] Then, if it is determined in step S19 that there has been silence for a predetermined period of time (step S19: YES), that is, if a predetermined period of time has elapsed since the sound acquisition unit 12 acquired sound (detected sound) but no sound has been acquired (no sound has been detected), the recording control unit 16 restores the quantization parameters of the image of interest to their original state (step S20). In other words, the setting of the image of interest is canceled.
[0125] Next, the process returns to step S10, and the subsequent steps are repeatedly executed.
[0126] As described above, according to Embodiment 2, since it is equipped with a sound direction detection unit 13, it is possible to detect the direction of arrival of sound around the vehicle, and since it is equipped with a focus image setting unit 14A, it is possible to dynamically set the image corresponding to the detected direction of arrival of the sound as the focus image. Furthermore, since it is equipped with an image encoding unit 15, it is possible to perform a first encoding process on the focus image among a plurality of images I, and a second encoding process on the images other than the focus image among the plurality of images I.
[0127] This allows for a clearer recording of the location of surrounding sounds (e.g., impact sounds, horns, sirens / brake sounds, screams / shouts) by highlighting the direction of arrival (the location of an incident such as an accident or a road rage incident) as a focus image, and enhancing the image quality compared to other images.
[0128] Using Figure 14, we will explain an example of the hardware configuration of the control device for the image recording device according to Embodiments 1 and 2.
[0129] Figure 14 shows an example of the hardware configuration of the control device for the image recording device according to Embodiments 1 and 2.
[0130] In Figure 14, the image recording device includes a processor 101 and a memory 102. The processor 101 may be, for example, a microprocessor, an MPU (Micro Processing Unit), or a CPU (Central Processing Unit). The processor 101 may include multiple processors. The memory 102 is composed of a combination of volatile memory and non-volatile memory. The memory 102 may include storage located away from the processor 101. In this case, the processor 101 may access the memory 102 via an I / O interface not shown.
[0131] Furthermore, each device in the above-described embodiment may be composed of hardware, software, or both, and may consist of one piece of hardware or software, or multiple pieces of hardware or software. The functions (processing) of each device in the above-described embodiment may be implemented by a computer. For example, a program for performing the operations in the embodiment may be stored in memory 102, and each function may be implemented by executing the program stored in memory 102 on the processor 101.
[0132] The programs described above can be stored and supplied to a computer using various types of non-temporary computer-readable media. Non-temporary computer-readable media include various types of tangible recording media. Examples of non-temporary computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memory (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). Programs may also be supplied to a computer using various types of temporary computer-readable media. Examples of temporary computer-readable media include electrical signals, optical signals, and electromagnetic waves. Temporary computer-readable media can be supplied to a computer via wired communication channels such as electric wires and optical fibers, or via wireless communication channels.
[0133] The present invention is not limited to the embodiments described above, and can be modified as appropriate without departing from the spirit of the invention. [Explanation of Symbols]
[0134] 1…Image recording device 1A…Image recording device 11…Image acquisition unit 12...Sound acquisition section 13...Sound direction detection unit 14, 14A, 14B... Area of focus setting section 15…Image encoding section 16…Recording Control Unit 17…Sound recognition unit 18…Records Department 19... Camera 20... Mike 101… Processor 102...memory AX…Optical axis I...Image P AX ...Optical axis position on Image I P ROI ...detection location ROI…area of interest V... Own vehicle
Claims
1. An image acquisition unit that acquires images of the area around the vehicle, A sound acquisition unit that acquires sounds from around the vehicle, A sound direction detection unit for detecting the direction from which the sound is coming, A focus area setting unit sets the region of the aforementioned image corresponding to the direction of arrival of the sound as the focus area, An image encoding unit performs a first encoding process on the region of interest in the image, and a second encoding process on the region of the image other than the region of interest. Equipped with, The image encoding unit, as the first encoding process, controls the quantization parameters so that the amount of code is greater than in the second encoding process if the volume of the acquired sound is less than a predetermined threshold, and controls the quantization parameters so that the amount of code is even greater than when the volume of the acquired sound is less than the threshold if the volume of the acquired sound is equal to or greater than a predetermined threshold. Image recording device.
2. An image acquisition unit that acquires multiple images of the area around the vehicle, A sound acquisition unit that acquires sounds from around the vehicle, A sound direction detection unit for detecting the direction from which the sound is coming, A focus image setting unit sets one or more images from the aforementioned plurality of images that correspond to the direction of arrival of the sound as focus images, An image encoding unit performs a first encoding process on the image of interest among the plurality of images, and a second encoding process on the images other than the image of interest among the plurality of images. Equipped with, The image encoding unit, as the first encoding process, controls the quantization parameters so that the amount of code is greater than in the second encoding process if the volume of the acquired sound is less than a predetermined threshold, and controls the quantization parameters so that the amount of code is even greater than when the volume of the acquired sound is less than the threshold if the volume of the acquired sound is equal to or greater than a predetermined threshold. Equipped with Image recording device.
3. Image acquisition step to acquire images of the area around the vehicle, Sound acquisition step to acquire sounds around the vehicle, A sound direction detection step for detecting the direction from which the sound is coming, A step of setting a region of interest, in which the region of the aforementioned image corresponding to the direction of arrival of the sound is set as the region of interest, An image encoding step in which a first encoding process is performed on the region of interest in the image, and a second encoding process is performed on the region of the image other than the region of interest, Equipped with, The image encoding step, as the first encoding process, performs a process that controls the quantization parameters so that the amount of encoding is greater than in the second encoding process if the volume of the acquired sound is less than a predetermined threshold, and performs a process that controls the quantization parameters so that the amount of encoding is even greater than when the volume of the acquired sound is less than the threshold. Image recording method.
4. Image acquisition step: Acquire multiple images of the area around the vehicle, Sound acquisition step to acquire sounds around the vehicle, A sound direction detection step for detecting the direction from which the sound is coming, A step of setting a focus image, in which one or more images from the plurality of images corresponding to the direction of arrival of the sound are set as focus images, An image encoding step in which a first encoding process is performed on the image of interest among the plurality of images, and a second encoding process is performed on the images other than the image of interest among the plurality of images, Equipped with, The image encoding step, as the first encoding process, performs a process that controls the quantization parameters so that the amount of encoding is greater than in the second encoding process if the volume of the acquired sound is less than a predetermined threshold, and performs a process that controls the quantization parameters so that the amount of encoding is even greater than when the volume of the acquired sound is less than the threshold. Image recording method.