Control system and device, imaging system and device, information processing device, control method, program, and storage medium

The control system enhances imaging devices by combining image and voice analysis to automatically capture images in lively environments, addressing the issue of missed shots in existing systems.

JP2025105015APending Publication Date: 2025-07-10CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023223263
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing imaging systems fail to recognize lively places or conversations, leading to missed shooting moments without user intervention.

Method used

A control system that combines image analysis and voice analysis to determine optimal shooting timing, using an imaging device with pan, tilt, and zoom functions, and a smartphone terminal for processing, to automatically capture images based on image and voice scores.

Benefits of technology

Suppresses missed shots by recognizing lively scenes or conversations, enabling automatic image capture without user operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025105015000001_ABST
    Figure 2025105015000001_ABST
Patent Text Reader

Abstract

To prevent missed shots without a user performing special operations on an imaging device.SOLUTION: Images repeatedly captured by an imaging apparatus are analyzed, and audio collected by sound collection means is also analyzed. On the basis of the obtained image analysis results and audio analysis results, it is determined whether to capture an image to be recorded. If it is determined that capturing should be performed, the imaging apparatus is instructed to capture the image for recording.SELECTED DRAWING: Figure 5A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a control system and device, an imaging system and device, an information processing device, a control method, a program, and a storage medium, and particularly to a technique for automatically controlling shooting timing.

Background Art

[0002] An imaging device that continuously takes pictures without the user giving a shooting instruction has been provided. As a technique related to such an imaging device, for example, Patent Document 1 discloses a technique related to an imaging device for acquiring a video of the user's preference without the user performing a special operation. Patent Document 1 discloses a technique of analyzing a captured image and performing shooting when it is determined that a subject currently exists within the shooting angle of view.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the shooting determination based only on the image analysis disclosed in Patent Document 1, it is impossible to recognize that the place is lively or that the conversation is lively, and there is a problem that the moment to be shot is missed.

[0005] The present invention has been made in view of the above problems, and an object thereof is to suppress misses without the user performing a special operation on the imaging device.

Means for Solving the Problems

[0006] To achieve the above object, a control system of the present invention for controlling the shooting timing of an image to be recorded by an imaging device includes an image analysis means for analyzing an image obtained by repeatedly shooting with the imaging device, a voice analysis means for analyzing the voice collected by a sound collection means, and based on the image analysis result obtained by the image analysis means and the voice analysis result obtained by the voice analysis means, it determines whether to shoot the image to be recorded, and when it is determined to shoot, it has a control means for instructing the imaging device to shoot the image to be recorded.

Effect of the Invention

[0007] According to the present invention, it is possible to suppress missed shots without the user performing a special operation on the imaging device.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 5C

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Embodiments for Carrying Out the Invention

[0009] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the invention according to the claims. Although a plurality of features are described in the embodiments, not all of these plurality of features are essential to the invention, and the plurality of features may be arbitrarily combined. Further, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.

[0010] FIG. 1 is a diagram showing the configuration of an imaging system including a control system in the present embodiment. The imaging system in the present embodiment includes an imaging device 101 and an information processing device 102.

[0011] In the present embodiment, the imaging device 101 is, for example, an automatic shooting camera that pans, tilts, zooms, and searches for a subject.

[0012] Further, in the present embodiment, a case where a smartphone terminal is used as the information processing device 102 will be described. Here, as an example of the information processing device 102, a smartphone terminal is used, but the information processing device 102 is not limited thereto, and may be, for example, a so-called tablet device, a personal computer, or the like. In the present embodiment, the functions provided by the information processing device 102 are realized in the form of an application operating on a smartphone terminal.

[0013] FIG. 2(a) is a schematic diagram showing the appearance of the imaging device 101, and FIG. 1(b) is a diagram for explaining the rotation direction.

[0014] The imaging device 101 shown in Fig. 2(a) is provided with an operation unit (not shown) such as buttons, switches, and touch panels for performing camera operations such as a power switch. The imaging unit 202 includes a lens barrel including a photographing lens group as an imaging optical system and an imaging element, and is attached to the fixing unit 203 via a pan rotation unit 205 which is a motor drive mechanism capable of rotating in the yaw direction (around the Y axis) shown in Fig. 2(b).

[0015] The tilt rotation unit 204 has a motor drive mechanism capable of rotating the imaging unit 202 in the pitch direction (around the X axis) shown in Fig. 2(b). By driving the pan rotation unit 205 and the tilt rotation unit 204, the orientation of the imaging optical system (i.e., the photographing direction) can be changed, and by controlling the driving of these rotation units, the orientation of the imaging optical system can be rotationally controlled in one or more directions.

[0016] Both the angular velocity meter 206 and the acceleration meter 207 are mounted on the fixing unit 203 of the imaging device 101. Then, based on the angular velocity meter 206 and the acceleration meter 207, the vibration of the imaging device 101 is detected, and the tilt rotation unit 204 and the pan rotation unit 205 are rotationally driven based on the detected shake angle. Thereby, the shake of the imaging unit 202 which is a movable part is corrected, or the inclination is corrected.

[0017] Fig. 3 is a block diagram showing the functional configuration of the imaging device 101 in the present embodiment. In Fig. 3, the control unit 320 is composed of a processor (for example, CPU, GPU, microprocessor, MPU, etc.) and a memory (for example, DRAM, SRAM, etc.). These execute various processes to control each block of the imaging device 101 and control data transfer between each block. The non-volatile memory (EEPROM) 312 is an electrically erasable and recordable memory, and stores constants, programs, etc. for the operation of the control unit 320.

[0018] The zoom unit 301 includes a zoom lens for performing zooming and is driven and controlled by a zoom drive control unit 302. The focus unit 303 includes a focus lens for performing focus adjustment and is driven and controlled by a focus drive control unit 304.

[0019] The imaging unit 306 includes an image sensor, receives light incident through each lens group, and converts it into electric charges according to the amount of the light. Then, the imaging unit 306 converts the obtained electric charges into an analog image signal, further performs A / D conversion, and outputs the obtained digital image data to an image processing unit 307. The image processing unit 307 applies image processing such as distortion correction, white balance adjustment, and color interpolation processing to the input digital image data and outputs the processed digital image data. The digital image data output from the image processing unit 307 is converted into a recording format such as the JPEG format by an image recording unit 308 and transmitted to a memory 311 and a video output unit 313.

[0020] The rotation drive unit 305 drives the tilt rotation unit 204 and the pan rotation unit 205 to rotationally drive the imaging unit 202 in the tilt direction and the pan direction.

[0021] As the device shake detection unit 319, for example, an angular velocity meter (gyro sensor) 206 that detects the angular velocity in the three-axis directions of the imaging device 101 and an accelerometer (acceleration sensor) 207 that detects the acceleration in the three-axis directions of the device are mounted. The device shake detection unit 319 calculates the rotation angle, shift amount, etc. of the imaging device 101 based on the detected signals.

[0022] The audio input unit 309 collects the sound around the imaging device 101 from the microphone provided in the imaging device 101, and transmits the digital audio signal obtained by analog-to-digital conversion to the audio processing unit 310. The audio processing unit 310 performs audio-related processing such as optimization processing on the input digital audio signal. Then, the digital audio signal processed by the audio processing unit 310 is transmitted to the memory 311 by the control unit 320. The memory 311 temporarily stores the image data and audio signal obtained by the image processing unit 307 and the audio processing unit 310 respectively.

[0023] Also, the image processing unit 307 and the audio processing unit 310 read out the image data and audio signal temporarily stored in the memory 311, perform encoding of the image data, encoding of the audio signal, etc., and generate a compressed image signal and a compressed audio signal. The control unit 320 transmits these compressed image signals and compressed audio signals to the recording and playback unit 316.

[0024] The recording and playback unit 316 records the compressed image signal, compressed audio signal, and other control data related to shooting generated by the image processing unit 307 and the audio processing unit 310 on the recording medium 317. When the audio signal is not compressed and encoded, the control unit 320 transmits the audio signal generated by the audio processing unit 310 and the compressed image signal generated by the image processing unit 307 to the recording and playback unit 316 for recording on the recording medium 317.

[0025] The recording medium 317 may be a recording medium built into the imaging device 101 or a removable recording medium. The recording medium 317 can record various data such as the compressed image signal, compressed audio signal, and audio signal generated by the imaging device 101, and a medium with a larger capacity than the non-volatile memory 312 is generally used. For example, the recording medium 317 includes all types of recording media such as hard disks, optical disks, magneto-optical disks, CD-Rs, DVD-Rs, magnetic tapes, non-volatile semiconductor memories, and flash memories.

[0026] In addition, the recording and playback unit 316 reads out the compressed image signal, compressed audio signal, audio signal, various data, and programs recorded on the recording medium 317. Then, the control unit 320 transmits the read compressed image signal, compressed audio signal, and audio signal to the image processing unit 307 and the audio processing unit 310, respectively. The image processing unit 307 and the audio processing unit 310 temporarily store the compressed image signal, compressed audio signal, and audio signal in the memory 311, decode them according to a predetermined procedure as necessary, and transmit the obtained signals to the video output unit 313 and the audio output unit 314, respectively.

[0027] The audio output unit 314 outputs, for example, a preset audio pattern or audio based on the audio signal transmitted from the audio processing unit 310 from a speaker built in the imaging device 101 during shooting or the like. Note that the audio output unit 314 may be an audio output terminal, and in this case, the audio signal is transmitted to a connected external speaker.

[0028] The LED control unit 315 controls, for example, the LED provided in the imaging device 101 during shooting or the like according to a preset lighting and blinking pattern.

[0029] The video output unit 313 is composed of, for example, a video output terminal, and transmits an image signal to display video on a connected external display or the like. Note that the audio output unit 314 and the video output unit 313 may be configured by a combined single terminal, for example, an HDMI (High-Definition Multimedia Interface) (registered trademark) terminal.

[0030] The communication unit 318 communicates between the imaging device 101 and the information processing device 102, and can, for example, transmit and receive audio signals, image signals, compressed audio signals, compressed image signals, and audio scores to be described later. In addition, a control signal related to shooting, such as a shooting start command, a shooting end command, panning, tilting, or zoom driving, is received from an external device that can communicate with the imaging device 101. The communication unit 318 is, for example, a wireless communication module such as an infrared communication module, a Bluetooth communication module, a wireless LAN communication module, WirelessUSB, or a GPS receiver.

[0031] Note that the voice input unit 309 in this embodiment has a configuration in which a plurality of microphones are mounted on the imaging device 101. The voice processing unit 310 can detect the direction of sound on the plane where the plurality of microphones are installed, and the obtained information is used for subject search and automatic shooting, which will be described later. Further, the voice processing unit 310 detects a trigger word. A trigger word is a predetermined word, and when the voice processing unit 310 recognizes it, it serves as a trigger for shooting. Details of the trigger word will be described later.

[0032] In addition, the voice processing unit 310 also performs sound scene recognition. In sound scene recognition, a neural network trained by machine learning based on a large amount of voice data in advance is used to determine the sound scene. For example, a neural network for detecting specific sound scenes such as "the cheers are rising", "clapping hands", and "making a sound" is set in the voice processing unit 310.

[0033] When a specific sound scene or a specific trigger word is detected, a detection trigger signal is output to the control unit 320.

[0034] As described above, the imaging device 101 in this embodiment is an automatic shooting camera. In automatic shooting, panning, tilting, and zooming are used to perform a subject search process, which will be described later, at a predetermined cycle, and the searched subject is automatically shot based on the image score and the voice score, which will be described later.

[0035] Here, the subject search process performed in the imaging device 101 will be described with reference to the flowchart of FIG. 10. Note that the subject search process is performed by the control unit 320.

[0036] First, when the subject search process is started, in S101, area division is performed for the entire area centered on the position of the imaging device 101.

[0037] Next, in S102, for each of the divided areas, an importance level indicating the priority of search is obtained according to the subject existing in each divided area and the situation of the scene of each divided area. The importance level based on the situation of the subject is calculated, for example, based on the number of subjects existing in each divided area, the size of the face, the orientation of the face, and the probability of face detection. Also, the importance level according to the situation of the scene of each divided area is obtained, for example, based on the general object recognition result, the scene discrimination result (blue sky, backlight, sunset scene, etc.), the sound level and voice recognition result from the direction of each divided area, the motion detection information within each divided area, and the like.

[0038] Next, in S103, a divided area whose obtained importance level is higher than a predetermined threshold is determined as a search target area. Then, in S104, the search target angles of pan and tilt necessary to capture the search target area as an angle of view are calculated.

[0039] Next, in S105, based on the calculated pan-tilt search target angles, the driving amounts of pan and tilt are calculated, and the tilt rotation unit 204 and the pan rotation unit 205 are driven. Also, in S106, it is determined whether a subject exists in the search target area. If a subject exists, in S107, the zoom driving amount is calculated based on the size of the subject in the subject recognition image, and the zoom unit 301 is driven by the zoom drive control unit 302. Through the above processing, the subject can be searched.

[0040] Although the method of performing subject search by driving the tilt rotation unit 204 and the pan rotation unit 205 and driving the zoom unit 301 has been described, the present invention is not limited to this. For example, a plurality of wide-angle lenses may be mounted on the imaging device 101, and the entire range may be photographed at once to perform subject search.

[0041] FIG. 4 is a block diagram showing the configuration of a smartphone terminal, which is an example of the information processing apparatus 102 in the present embodiment.

[0042] The control unit 401 controls each part of the information processing apparatus 102 according to the input signals and the programs described later. Note that instead of the control unit 401 controlling the whole, a plurality of hardware may share the processing to control the entire apparatus.

[0043] The imaging unit 402 converts subject light imaged by a lens included in the imaging unit 402 into an electrical signal, and outputs digital data obtained by performing noise reduction processing and the like as image data. The captured image data is stored in a buffer memory included in the working memory 404, and then predetermined calculations are performed by the control unit 401 and recorded on the recording medium 407.

[0044] The non-volatile memory 403 is an electrically erasable and recordable non-volatile memory, and stores an OS (operating system) which is basic software executed by the control unit 401, various programs, and the like. A program for communicating with the imaging apparatus 101 is also held in the non-volatile memory 403 and is assumed to be installed as a communication application. Note that the processing of the information processing apparatus 102 in the present embodiment is realized by reading a program provided by the communication application. Note that the communication application has a program for using basic functions (for example, a wireless LAN function or a Bluetooth function) of the OS installed in the information processing apparatus 102. Note that the OS of the information processing apparatus 102 may have a program for realizing the processing in the present embodiment.

[0045] The working memory 404 is used as a buffer memory for temporarily storing image data generated by the imaging unit 402 and data received from the imaging apparatus 101, an image display memory for the display unit 406, a working area of the control unit 401, and the like.

[0046] The operation unit 405 is used to receive instructions from the user for the information processing apparatus 102. The operation unit 405 includes, for example, a power button for the user to instruct the ON / OFF of the power supply of the information processing apparatus 102, and operation members such as a touch panel formed on the display unit 406.

[0047] The display unit 406 performs display of image data, character display for interactive operations, and the like. Note that the display unit 406 does not necessarily have to be mounted on the information processing apparatus 102. The information processing apparatus 102 may be connected to the display unit 406 and only needs to have at least a display control function for controlling the display of the display unit 406.

[0048] The recording medium 407 can record the image data output from the imaging unit 402 and the image data received from the communication device. The recording medium 407 may be configured to be detachable from the information processing apparatus 102 or may be built into the information processing apparatus 102. That is, the information processing apparatus 102 only needs to have means for accessing at least the recording medium 407.

[0049] The connection unit 408 is an interface for connecting to the imaging apparatus 101. The information processing apparatus 102 in the present embodiment can exchange data with the imaging apparatus 101 via the connection unit 408. In the present embodiment, the connection unit 408 includes an interface for communicating with the imaging apparatus 101 via a wireless LAN. The control unit 401 realizes wireless communication with the imaging apparatus 101 by controlling the connection unit 408.

[0050] The public network connection unit 409 is an interface used when performing public wireless communication. The information processing device 102 can make calls and perform data communication with other devices via the public network connection unit 409. During a call, the control unit 401 inputs and outputs voice signals via the microphone 410 and the speaker 411. In this embodiment, the public network connection unit 409 includes an interface for performing communication using 4G. Note that not limited to 4G, other communication methods such as LTE, WiMAX, ADSL, FTTH, and even 5G may be used. Also, the connection unit 408 and the public network connection unit 409 do not necessarily have to be configured with independent hardware, and for example, they can be shared by one antenna.

[0051] The short-range wireless communication unit 412 is an interface for short-range wireless connection with other communication devices. The information processing device 102 in this embodiment can exchange data with the imaging device 101 via the short-range wireless communication unit 412.

[0052] Next, the overall flow of the processing performed by the imaging system in this embodiment will be described with reference to FIGS. 5A to 5C. In the imaging system in this embodiment, as described above, the control unit 320 of the imaging device 101 performs subject search processing at a predetermined period. And the imaging system in this embodiment performs a series of processes related to shooting shown in FIG. 5A in parallel with the subject search processing.

[0053] In the imaging system in this embodiment, the imaging device 101 generates a subject recognition image at a predetermined period, and obtains an image score obtained by quantifying the result of analyzing this subject recognition image. In parallel, the imaging device 101 constantly collects sound and transmits voice data to the information processing device 102. Further, the information processing device 102 analyzes the voice data received from the imaging device 101 at a predetermined period, and transmits a voice score obtained by quantifying the voice analysis result to the imaging device 101. The imaging device 101 determines whether to perform shooting based on the image score and the voice score, and controls the shooting timing.

[0054] The flow of the process described above will be described in detail with reference to FIG. 5A. In S501, the control unit 320 of the imaging device 101 collects sound. Next, in S502, the control unit 320 of the imaging device 101 transmits the collected audio data to the information processing device 102 via the communication unit 318 of the imaging device 101. At this time, in order to accurately identify the order of the audio data in subsequent processing, a timestamp is added to the audio data and then transmitted. Next, in S503, the control unit 401 of the information processing device 102 stores the received audio data in the working memory 404. At this time, the audio data is stored in the data format of a queue that can identify the order.

[0055] The imaging device 101 and the information processing device 102 repeatedly execute S501 to S503 described above at a predetermined cycle.

[0056] Next, in S504, the control unit 401 of the information processing device 102 sequentially acquires the audio data stored in the queue format in S503 from the head. At this time, referring to the timestamp added to the audio data, the audio data for which a predetermined time or more has elapsed from the current time is deleted after acquisition. Next, in S505, the control unit 401 of the information processing device 102 performs speech recognition using the audio data acquired in S504. Details of the speech recognition will be described later.

[0057] Next, in S506, the control unit 401 of the information processing device 102 performs speech analysis. In the speech analysis, a speech score is calculated based on the result of the speech recognition in S505. Details of the speech score will be described later. Next, in S507, the control unit 401 of the information processing device transmits the speech score obtained in S506 to the imaging device 101 via the connection unit 408 of the information processing device 102. Here too, in order to accurately identify the order of the speech scores in subsequent processing, the same timestamp as the timestamp added to the audio data in S504 is added to the speech score and then transmitted. Next, in S508, the control unit 320 of the imaging device 101 stores the received speech score in the memory 311.

[0058] The imaging device 101 and the information processing device 102 repeatedly execute the series of processes S504 to S508 described above at a predetermined cycle.

[0059] On the other hand, in S509, the imaging unit 306 of the imaging device 101 performs imaging. Here, the imaging process in this embodiment will be described. First, the imaging unit 306 captures images at a predetermined cycle to generate a subject recognition image. The subject recognition image generated here is used for the resident subject search process and the image analysis in S510 described later. Next, the generated subject recognition image is stored in the memory 311.

[0060] Next, in S510, the control unit 320 of the imaging device 101 performs image analysis. In the image analysis, the image captured in S509 is analyzed to calculate an image score. The image score is calculated using the number of human faces, the smile degree of the face, the eye closure degree, the face position, the face angle, and the gaze angle of the subject in the current subject recognition image. After the calculation of the image score is completed, the subject recognition image is deleted from the memory 311. In addition to the subject detection results described above, animal detection results, general object recognition results, scene discrimination results, etc. may be further used for calculation.

[0061] Next, in S511, the control unit 320 of the imaging device determines whether to output the captured image. Here, it is determined whether the total value of the voice score stored in S508 and the image score calculated in S510 exceeds a threshold Th1. As a result of the determination, if the total value exceeds the threshold Th1, the process proceeds to S512. On the other hand, if it does not exceed the threshold Th1, this process ends. Note that the image score and the voice score in this embodiment are both normalized values with a minimum value of 0 and a maximum value of 100 so that they are equivalent values and can be added.

[0062] When it is determined in S511 that the threshold Th1 is exceeded, in S512, the captured image is output. The captured image output here means performing a shooting operation to generate a captured image and storing it in a storage medium, which is different from the imaging in S509.

[0063] The imaging system in this embodiment performs the processes of S501 to S512 described above in parallel.

[0064] Although the method of determining whether the total value of the image score and the voice score exceeds the threshold Th1 in S511 has been described, the present invention is not limited to this. For example, a method may be used in which it is determined whether the image score and the voice score each exceed a different threshold, and it is determined whether at least one of them exceeds the threshold.

[0065] Also, the sound collection process may be performed by the information processing device 102 instead of the imaging device 101. However, in that case, the sound collected may not be the sound around the imaging device 101. Therefore, in order to capture the person who spoke within the shooting angle, it is applicable when the distance between the imaging device 101 and the information processing device 102 is within a predetermined distance. Also, sound may be collected by a microphone connected to the information processing device 102 by wire or wirelessly.

[0066] In the process shown in FIG. 5A, the imaging determination of whether to output the captured image in S511 is made by the imaging device 101. However, as shown in FIG. 5B, it may be configured to be made by the information processing device 102. In FIG. 5B, the voice score is not transmitted to the imaging device 101, and the image score, which is the result of image analysis in S510, is transmitted from the imaging device 101 to the information processing device 102 and stored. Then, in S511, the information processing device 102 makes a shooting determination based on the image score and the voice score. When shooting is to be performed, shooting is performed by outputting a shooting instruction to the imaging device 101 in S515. In this way, the shooting timing is controlled based on the image score and the voice score.

[0067] Also, in the process in FIG. 5A, the voice recognition in S505 and the voice analysis in S506 are configured to be performed by the information processing device 102. However, as shown in FIG. 5C, it may be configured to be performed by the imaging device 101.

[0068] In this case, it is necessary to let the voice processing unit 310 of the imaging device 101 learn voice patterns of several words based on a large amount of voice data in advance. Therefore, in the configuration where the voice recognition in S505 is performed by the imaging device 101, it is applicable when the pre-learned voice pattern can be installed in the voice processing unit 310 of the imaging device 101. On the other hand, in the configuration where the voice recognition in S505 is performed by the information processing device 102, or in the configuration where the information processing device 102 transmits voice data to an external server, software, cloud service, etc. for voice recognition and receives the voice recognition result, it is possible to perform highly accurate analysis using the ever-developing voice recognition technology.

[0069] Hereinafter, in this embodiment, the method shown in FIG. 5A will be described.

[0070] Next, the processes S505 and S506 performed by the information processing device 102 in this embodiment will be described in detail.

[0071] There are three methods for calculating the voice score. The first method is by detecting a human voice, the second method is by detecting a topic, and the third method is by detecting a trigger word registered by the user. The specific processes for each will be described later. The voice score in this embodiment is a value obtained by summing up three bonus scores calculated by these three methods, and this total value is a value normalized in a value range with a minimum value of 0 and a maximum value of 100. By determining based on the total value of the three bonus scores, it is possible to perform highly accurate scene detection compared to the case of determining the bonus scores of each method separately. Note that in this embodiment, when the total value of the three bonus scores exceeds 100, the voice score is normalized to 100.

[0072] In addition, the control unit 401 of the information processing apparatus performs the calculation processes of the three scores at predetermined time intervals. In the present embodiment, for example, the first method is performed at intervals of 2 seconds, the second method is performed at intervals of 10 seconds, and the third method is performed at intervals of 2 seconds. That is, in the present embodiment, the calculation process of summing and normalizing the score points obtained by the first method and the third method is performed at intervals of 2 seconds, and the calculation process of summing and normalizing the score points obtained by all three methods is performed at intervals of 10 seconds.

[0073] First, the voice score calculation method by detecting human voices, which is the first method, will be described with reference to FIG. 6. This method utilizes the fact that when the input voice is a human voice or the like, it is recognized as voice, but when it is noise such as the sound of a vacuum cleaner or footsteps, it is not recognized as voice. This is because when voice is recognized, it can be inferred that people are having a conversation and it is likely to be the timing to take a photo. The control unit 401 of the information processing apparatus 102 in the present embodiment executes the following processes at a predetermined time interval (for example, at intervals of 2 seconds).

[0074] First, in S601, voice recognition is performed. An example of the flow of the voice recognition process will be described below, but the voice recognition method is not limited to this. First, feature amounts such as the frequency and intensity of the voice are extracted and converted into quantitative numerical values. Next, based on a learning pattern learned in advance from a large amount of voice data, the phoneme closest to the extracted feature amount is extracted. A phoneme is the smallest unit of sound. The number and types of phonemes vary depending on the language. For example, in English, they are vowels and consonants, but in Japanese, in addition to vowels and consonants, there are nasal sounds, stop sounds, and long sounds. Next, pattern matching is performed using dictionary data in which pronunciations and words are registered, and the phonemes are converted into words. Voice recognition is performed through the above flow.

[0075] Note that the voice recognized here is not limited to Japanese or English, and may be other languages. Also, the voice recognized here is not limited to human voices, and may be voices of a predetermined type, for example, animal calls. However, the second and third methods described later assume that the voice is a human voice. Therefore, only when the voice recognized by the first method is an animal call, the voice score is not the sum of the three bonus scores, but a value obtained by multiplying the bonus score calculated by the first method by a predetermined value. Note that it is not limited to this, and it may be a value obtained by adding a predetermined value to the bonus score calculated by the first method.

[0076] An example of the flow of the process for determining animal calls will be described below, but the determination method is not limited to this. First, calculate the frequency spectrum of the collected voice and convert it into a voiceprint image. The frequency spectrum is the result of performing a Fourier transform on the voice and represents the frequency components of the voice. Next, based on a learning pattern learned in advance from a large amount of animal voiceprint image data, determine whether the converted voiceprint image is an animal call. Through the above process, animal calls are recognized in voice recognition.

[0077] Next, in S602, it is determined whether voice recognition was successful in S601. If voice recognition was successful, the process proceeds to S603. If voice recognition was not successful, this process ends without adding points to the voice score.

[0078] Next, in S603, it is determined whether the volume of the voice exceeds the threshold Th2. If it exceeds the threshold Th2, the process proceeds to S604. If it does not exceed the threshold Th2, this process ends without adding points to the voice score.

[0079] Then, when voice recognition is successful in S602 and the volume of the voice exceeds the threshold Th2 in S603, points are added to the voice score in S604. In this embodiment, the bonus scores are determined in four levels: +25 points, +50 points, +75 points, and +100 points, such that the higher the volume of the voice, the higher the bonus score.

[0080] By the process described above, when voice recognition is possible and the volume of the voice exceeds a predetermined threshold, a voice score is added according to the volume of the voice. This enables the voice score to be added only when the atmosphere is enlivened by a human voice, without the voice score being added due to loud noises or ambient sounds such as the sound of a vacuum cleaner or footsteps.

[0081] In this embodiment, the case of performing voice recognition within the information processing apparatus 102 has been described. However, the communication unit may transmit voice data to an external server, software, cloud service, etc. for voice recognition and receive the voice recognition result.

[0082] Next, a method for calculating a voice score by detecting a topic, which is the second method, will be described with reference to FIG. 7. This method utilizes the fact that when the same word is detected a predetermined number of times or more in the text of the voice detected within a predetermined time, it can be inferred that the conversation is lively on the topic related to that word. The control unit 401 of the information processing apparatus 102 in this embodiment executes the following processing at predetermined time intervals (for example, at 10 - second intervals).

[0083] First, in S701, voice recognition is performed. Since the flow of the voice recognition process is the same as that of S601, the description thereof is omitted.

[0084] Next, in S702, it is determined whether voice recognition was successful in S701. If voice recognition was successful, the process proceeds to S703. If voice recognition was not successful, this process ends without adding a voice score.

[0085] Next, in S703, morphological analysis is performed using the text sentence obtained by voice recognition. A morpheme is the smallest unit of a word with meaning, and morphological analysis is to divide a sentence into morphemes and discriminate part-of-speech information such as nouns and verbs. For example, in the case of the sentence "I eat a red apple.", it is divided into "I", "eat", "a", "red", "apple", where "I" is a pronoun, "eat" is a verb, "a" is an article, "red" is an adjective, and "apple" is a noun, and each is discriminated accordingly. Note that in this embodiment, morphological analysis is exemplified using a Japanese text sentence, but it is not limited to Japanese. For example, in the case of an English sentence "I eat an red apple.", it is divided into "I", "eat", "an", "red", "apple", where "I" is a pronoun, "eat" is a verb, "an" is an article, "red" is an adjective, and "apple" is a noun, and each is discriminated accordingly. Thus, morphological analysis is possible even for sentences other than Japanese.

[0086] Next, in S704, it is determined whether any same morpheme has been detected a predetermined number of times within a predetermined time. The predetermined time in this embodiment is 10 seconds as described above. The morpheme to be detected here is the morpheme discriminated as a noun in S703. For example, in the case of the sentence "Dogs are cute. Big dogs and small dogs are both cute.", it can be determined that the same morpheme "dog" has been detected 3 times.

[0087] Thus, when any same morpheme has been detected a predetermined number of times within a predetermined time in S704, the process proceeds to S705, and when not detected, this process ends without adding points to the voice score.

[0088] And when any same morpheme has been detected more than a predetermined number of times within a predetermined time in S704, it is determined that the conversation is lively on the topic related to that morpheme, and in S705, points are added to the voice score. In this embodiment, the number of points added increases as the number of detections increases. When the number of detections is 2 times, +25 points are added, and when it is 3 times or more, +50 points are added.

[0089] By the process described above, it is possible to detect that there is a buzz about a certain topic.

[0090] In the present embodiment, the case where morphological analysis is performed within the information processing apparatus 102 has been described. However, the communication unit may transmit the voice data to an external server, software, cloud service, etc. for morphological analysis and receive the morphological analysis result.

[0091] Next, a third method, a method of detecting a trigger word registered by the user and calculating a voice score, will be described with reference to FIGS. 8 and 9.

[0092] FIG. 8 is a screen for registering a trigger word, which is displayed on the display unit 406 of the information processing apparatus 102 in the present embodiment. The constituent members of the screen will be described below. First, 801 is a trigger word. When the registered trigger word is recognized, a voice score is added. A trigger word is one that registers words that are often uttered at the timing when the user wants to take a picture. This makes it possible to take pictures using natural daily conversations as a trigger without the user being conscious of taking pictures.

[0093] For example, if the word "cute" exemplified as the first of the trigger words 801 is registered, it can be pre-registered to take a picture in response to the voice of "cute" uttered by the parent at the moment when the child makes cute actions.

[0094] Also, if "Taro" exemplified as the name of the child in the second of the trigger words 801 is registered, it can be pre-registered to take a picture at the moment when the parent calls the child's name "Taro". In this way, the trigger word can be freely registered by the user.

[0095] 802 is an edit button. When pressed, the trigger word registered once can be edited. 803 is a delete button. When pressed, the trigger word registered once can be deleted. And 804 is an additional button. When pressed, a trigger word can be newly added and registered.

[0096] FIG. 9 is a diagram showing the flow of a process for detecting a trigger word registered by a user and calculating an audio score. The control unit 401 of the information processing apparatus 102 in the present embodiment executes the following processes at a predetermined time interval (for example, at intervals of 2 seconds).

[0097] First, in S901, speech recognition is performed. Since the flow of the speech recognition process is the same as that of S601, the description thereof is omitted.

[0098] Next, in S902, it is determined whether speech recognition was successful in S901. If speech recognition was successful, the process proceeds to step S903. If speech recognition was not successful, this process ends without adding points to the audio score.

[0099] In S903, it is determined whether a trigger word pre-registered by the user was detected from the text sentence obtained by speech recognition. If detected, the process proceeds to S904. If not detected, this process ends without adding points to the audio score.

[0100] And if a trigger word is detected in S903, points are added to the audio score in S904. In the present embodiment, for example, the number of additional points when a trigger word is detected is set to +100 points.

[0101] By the method described above, a user can issue a shooting instruction to the imaging device using a trigger word freely registered by the user.

[0102] With the process flow described above, the imaging system in the present embodiment can, in addition to image analysis, also use audio analysis in combination, enabling recognition of a lively scene or a lively conversation that could not be achieved by image analysis alone. As a result, it is possible to suppress missed shots.

[0103] Note that the imaging system in this embodiment is configured such that the imaging device is provided with pan, tilt, and zoom functions, whereby autonomous search of a subject person and automatic adjustment of the imaging angle can be performed. On the other hand, the present invention is also applicable to an imaging device that does not have components for realizing such pan, tilt, and zoom functions. For example, even in an imaging system using an imaging device that captures images with a fixed imaging angle or zoom ratio, such as a surveillance camera, automatic shooting using a combination of image analysis and voice analysis is useful, and the present invention can be applied.

[0104] In addition, in the automatic shooting in this embodiment, although it has been described that shooting is automatically performed based on the image score and the voice score, the shooting conditions are not limited to this. For example, the current zoom ratio, the elapsed time since the previous shooting, the shooting time, etc. may be further used.

[0105] In addition, although it has been described that the voice score in this embodiment is a value obtained by summing three scores calculated by three types of calculation methods, it is not necessary to perform all detections by the three methods, and one or more of them may be combined and used. In that case, the point values calculated in each method and the method for calculating the voice score using those values are not limited to the content described above, and may be changed as appropriate.

[0106] <Other Embodiments> Note that the present invention may be applied to a system composed of a plurality of devices or to an apparatus composed of a single device.

[0107] In addition, the present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiment to a system or apparatus via a network or a storage medium, and having one or more processors in the computer of the system or apparatus read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.

[0108] <Summary> The disclosure of this embodiment includes the following configurations.

[0109] (Item 1) A control system for controlling the shooting timing of an image to be recorded by an imaging device, comprising: image analysis means for analyzing an image obtained by repeatedly shooting with the imaging device; voice analysis means for analyzing the voice collected by the voice collection means; Based on the image analysis result obtained by the image analysis means and the voice analysis result obtained by the voice analysis means, it is determined whether to shoot the image to be recorded. When it is determined to shoot, the imaging device is controlled to instruct shooting of the image to be recorded A control system characterized by comprising the above. (Item 2) The voice analysis means according to item 1, wherein the voice analysis means analyzes the characteristics of the voice and outputs a voice score obtained by quantifying the voice. (Item 3) The voice analysis means according to item 2, wherein the voice analysis means detects a predetermined type of voice from the voice, and when the magnitude of the detected predetermined type of voice is greater than a predetermined threshold value, points are added to the voice score. (Item 4) The voice analysis means according to item 3, wherein the greater the magnitude of the predetermined type of voice, the greater the value added to the voice score. (Item 5) The control system according to item 3 or 4, wherein the predetermined voice is a human voice. (Item 6) The voice analysis means according to any one of items 2 to 5, wherein when the same morpheme is repeatedly detected from the voice more than a predetermined number of times within a predetermined time, points are added to the voice score. (Item 7) The control system according to item 6, wherein when the number of morphemes repeatedly detected within a predetermined time by the voice analysis means is the first number, a larger value is added to the voice score than when it is the second number which is less than the first number. (Item 8) The control system according to any one of items 2 to 7, wherein the voice analysis means adds points to the voice score when a predetermined word is detected from the voice. (Item 9) An imaging system including the control system according to any one of items 1 to 8, an imaging device having an imaging means for taking an image, the image analysis means, and the control means, a control device including the voice analysis means, and a communication means for communicating between the imaging device and the control device characterized by comprising the same. (Item 10) A control device for controlling the imaging timing of an image to be recorded by an imaging device, an image analysis means for analyzing an image obtained by repeatedly taking images by the imaging device, a voice analysis means for analyzing a voice collected by a sound collection means, a control means for determining whether or not to take an image to be recorded based on the image analysis result obtained by the image analysis means and the voice analysis result obtained by the voice analysis means, and when it is determined to take an image, instructing the imaging device to take an image to be recorded characterized by comprising the same. (Item 11) The control device according to item 10, wherein the voice analysis means outputs a voice score obtained by analyzing and digitizing the characteristics of the voice. (Item 12) The control device according to item 11, wherein the voice analysis means detects a predetermined type of voice from the voice, and adds points to the voice score when the magnitude of the detected predetermined type of voice is greater than a predetermined threshold value. (Item 13) The control device according to item 12, wherein the voice analysis means adds a larger value to the voice score as the magnitude of the voice of the predetermined type is larger. (Item 14) The control device according to item 12 or 13, wherein the predetermined voice is a human voice. (Item 15) The control device according to any one of items 11 to 14, wherein the voice analysis means adds a point to the voice score when the same morpheme is repeated a predetermined number of times or more within a predetermined time from the voice. (Item 16) The control device according to item 15, wherein the voice analysis means adds a larger value to the voice score when the number of times of the morpheme repeated within a predetermined time is the first number of times than when it is the second number of times which is less than the first number of times. (Item 17) The control device according to any one of items 11 to 16, wherein the voice analysis means adds a point to the voice score when a predetermined word is detected from the voice. (Item 18) A control device according to any one of items 10 to 17, imaging means for taking an image, and An imaging device characterized by comprising the same. (Item 19) A control device according to any one of items 10 to 17, communication means for communicating with the imaging device, and An information processing device characterized by comprising the same.

[0110] A control method for controlling the imaging timing of an image to be recorded by an imaging device, an image analysis step of analyzing an image obtained by repeatedly imaging by the imaging device, a voice analysis step of analyzing a voice collected by a sound collection means, and Based on the image analysis result obtained in the image analysis step and the voice analysis result obtained in the voice analysis step, it is determined whether to capture an image to be recorded. When it is determined to capture, a control step for instructing the imaging device to capture the image to be recorded, and A control method characterized by comprising the above. (Item 21) A program for causing a computer to function as each means of the control system according to any one of Items 1 to 8. (Item 22) A program for causing a computer to function as each means of the control device according to any one of Items 10 to 17. (Item 23) A computer-readable storage medium storing the program according to Item 21 or 22.

[0111] The invention is not limited to the above embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Therefore, claims are attached to disclose the scope of the invention.

Explanation of Reference Numerals

[0112] 101: Imaging device, 102: Information processing device, 202: Imaging unit, 203: Fixing part, 204: Tilt rotation unit, 205: Pan rotation unit, 206: Angular velocity meter, 207: Accelerometer, 301: Zoom unit, 302: Zoom drive control unit, 303: Focus unit, 304: Focus drive control unit, 305: Rotation drive unit, 306: Imaging unit, 307: Image processing unit, 308: Image recording unit, 309: Voice input unit, 310: Voice processing unit, 311: Memory, 312: Non-volatile memory, 313: Video output unit, 314: Voice output unit, 315: LED control unit, 316: Recording / playback unit, 317: Recording medium, 318: Communication unit, 319: Device shake detection unit, 320: Control unit, 401: Control unit, 403: Non-volatile memory, 404: Working memory, 405: Operation unit, 406: Display unit, 407: Storage medium, 408: Connection unit, 412: Short-range wireless communication unit

Claims

1. A control system for controlling the shooting timing of an image to be recorded by an imaging device, comprising: image analysis means for analyzing an image obtained by repeatedly shooting with the imaging device; voice analysis means for analyzing the voice collected by the voice collection means; Based on the image analysis result obtained by the image analysis means and the voice analysis result obtained by the voice analysis means, it is determined whether to shoot the image to be recorded. When it is determined to shoot, the imaging device is instructed to shoot the image to be recorded. control means; A control system characterized by comprising.

2. The control system according to claim 1, wherein the voice analysis means outputs a voice score obtained by analyzing and digitizing the characteristics of the voice.

3. The control system according to claim 2, wherein the voice analysis means detects a predetermined type of voice from the voice, and adds points to the voice score when the magnitude of the detected predetermined type of voice is greater than a predetermined threshold value.

4. The control system according to claim 3, wherein the voice analysis means adds a larger value to the voice score as the magnitude of the predetermined type of voice is larger.

5. The control system according to claim 3, wherein the predetermined voice is a human voice.

6. The control system according to claim 2, wherein the voice analysis means adds points to the voice score when the same morpheme is repeatedly detected from the voice more than a predetermined number of times within a predetermined time.

7. The control system according to claim 6, wherein when the number of times of the morpheme repeatedly detected within a predetermined time is the first number of times, the voice analysis means adds a larger value to the voice score than when it is the second number of times less than the first number of times.

8. The control system according to claim 2, wherein the voice analysis means adds points to the voice score when a predetermined word is detected from the voice.

9. An imaging system including the control system according to any one of claims 1 to 8, comprising: an imaging means for shooting an image, the imaging device having the image analysis means and the control means; a control device including the voice analysis means; communication means for communicating between the imaging device and the control device An imaging system characterized by having **Claim 10** A control device for controlling the imaging timing of an image to be recorded by an imaging device, image analysis means for analyzing an image obtained by repeatedly imaging with the imaging device; voice analysis means for analyzing the voice collected by the voice collection means; Based on the image analysis result obtained by the image analysis means and the voice analysis result obtained by the voice analysis means, it is determined whether to capture an image to be recorded. When it is determined to capture, the imaging device is instructed to capture the image to be recorded. control means A control device characterized by having **Claim 11** The control device according to claim 10, wherein the voice analysis means outputs a voice score obtained by analyzing and digitizing the characteristics of the voice. **Claim 12** The control device according to claim 11, wherein the voice analysis means detects a predetermined type of voice from the voice, and adds points to the voice score when the magnitude of the detected predetermined type of voice is greater than a predetermined threshold. **Claim 13** The control device according to claim 12, wherein the voice analysis means adds a larger value to the voice score as the magnitude of the predetermined type of voice is larger. **Claim 14** The control device according to claim 12, wherein the predetermined voice is a human voice. **Claim 15** The control device according to claim 11, wherein the voice analysis means adds points to the voice score when the same morpheme is repeated a predetermined number of times or more within a predetermined time from the voice. **Claim 16** The control device according to claim 15, wherein when the number of times of the morpheme repeated within a predetermined time is the first number of times, the voice analysis means adds a larger value to the voice score than when it is the second number of times less than the first number of times. **Claim 17** The control device according to any one of claims 11 to 16, wherein the voice analysis means adds points to the voice score when a predetermined word is detected from the voice. **Claim 18** The control device according to any one of claims 10 to 17, imaging means for capturing an image An imaging device characterized by having **Claim 19** The control device according to any one of claims 10 to 17, communication means for communicating with the imaging device An information processing apparatus characterized by having

20. A control method for controlling the imaging timing of an image to be recorded by an imaging device, comprising: an image analysis step of analyzing an image obtained by repeatedly imaging with the imaging device; a voice analysis step of analyzing voice collected by a voice collection means; a control step of determining whether to capture an image to be recorded based on the image analysis result obtained in the image analysis step and the voice analysis result obtained in the voice analysis step, and instructing the imaging device to capture the image to be recorded when it is determined to capture; A control method characterized by having

21. A program for causing a computer to function as each means of the control system according to any one of Claims 1 to 8.

22. A program for causing a computer to function as each means of the control device according to any one of Claims 10 to 17.

23. A computer-readable storage medium storing the program according to Claim 21.

24. A computer-readable storage medium storing the program according to Claim 22.

Citation Information

Patent Citations

  • Imaging device and control method thereof

    JP6766086B2