Information processing device, method for controlling information processing device, and program
The information processing device addresses the challenge of creating layouts for desired scenes by acquiring image and audio data, converting audio to text, and generating layouts with associated text, effectively preserving private moments.
Patent Information
- Application Number
- JP2024014878
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-02
- Publication Date
- 2025-08-15
AI Technical Summary
Existing technologies face challenges in selecting and creating a layout for desired scenes, such as private moments, where it is difficult to capture and preserve specific images with associated audio data.
An information processing device that acquires image and audio data, converts audio to text, selects desired images, and generates a layout image with associated text, using technologies like deep learning for voice recognition and image analysis.
Enables the selection and creation of a layout image that includes desired images and their associated text, allowing for effective preservation of private moments.
Smart Images

Figure 2025119830000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, a control method for an information processing device, and a program. [Background technology]
[0002] Speech recognition technology has been known for some time as a technique for recognizing and converting audio data containing recorded human conversations into text. Patent Document 1 describes a minutes-generating device that converts speech generated during a meeting into text using speech recognition technology and then generates a summary text summarizing the content. The minutes-generating device described in Patent Document 1 can create a layout by associating the summary text with image data used during the meeting. This layout serves to leave evidence of the meeting. Also known is a camera that pre-stores personal information, such as a facial photograph. This camera can track a subject based on the personal information and capture photos or videos of the subject at desired times. When using such a camera to capture private moments, such as childcare, a layout can be created by combining image data and audio data recorded by the camera. However, this is different from the layout creation for a meeting. Specifically, for a meeting, it is generally preferable to create a layout covering the entire meeting from start to finish. However, for private moments, it is preferable to select a scene that is relatively clear and that you want to keep as a memory, and create a layout for that scene. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-149083 Summary of the Invention [Problem to be solved by the invention]
[0004] However, in Patent Document 1, there is a problem that it is difficult to select a desired scene that one wants to keep as a memory and create a layout of that scene.
[0005] The present invention has been made in view of the above-mentioned problems, and aims to provide an information processing device, a control method for the information processing device, and a program that can select a desired image and acquire a layout image that includes the selected image and text associated with the image. [Means for solving the problem]
[0006] In order to achieve the above object, the information processing device of the present invention is characterized by comprising: an acquisition means for acquiring image data including an image of a person and audio data associated with the image data; a conversion means for converting the audio data acquired by the acquisition means into text data; a selection means for selecting the image data acquired by the acquisition means; and a generation means for generating a layout image in which specific text data of the audio uttered by the person contained in the image data selected by the selection means and the image data selected by the selection means are arranged within the text data. [Effects of the Invention]
[0007] According to the present invention, it is possible to select a desired image and obtain a layout image that includes the selected image and text associated with the image. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 2 is a block diagram showing a hardware configuration of the layout image generating system. [Figure 2] FIG. 2 is a block diagram showing the hardware configuration of the imaging device. [Figure 3] FIG. 2 is a block diagram showing the software configuration of the imaging device. [Figure 4] FIG. 2 is a block diagram showing a hardware configuration of the information processing device. [Figure 5] FIG. 2 is a block diagram showing a software configuration of the information processing device. [Figure 6] FIG. 2 is a block diagram illustrating a hardware configuration of the printing apparatus. [Figure 7] FIG. 2 is a block diagram illustrating a software configuration of the printing apparatus. [Figure 8] 10 is a flowchart showing a process executed by the imaging device. [Figure 9] FIG. 2 is a diagram showing an example of person information stored in the image capturing device. [Figure 10A] 10 is a flowchart showing a process executed by the information processing device. [Figure 10B] FIG. 1 is an image diagram illustrating an example of processing executed by an information processing device. [Figure 11] 10 is a flowchart illustrating a process executed by the printing device. [Figure 12] FIG. 10 is a diagram illustrating an example of a layout image. [Figure 13] 10B is a flowchart showing detailed processing (text conversion processing) in step S1001, which is a subroutine of the flowchart shown in FIG. 10A. [Figure 14] FIG. 10 is a diagram illustrating an example of text information. [Figure 15] 10B is a flowchart showing detailed processing (face extraction processing) in step S1003, which is a subroutine of the flowchart shown in FIG. 10A. [Figure 16] FIG. 10 is a diagram illustrating an example of face area information. [Figure 17] 10B is a flowchart showing detailed processing (layout image generation processing) in step S1004, which is a subroutine of the flowchart shown in FIG. 10A. [Figure 18A] FIG. 10 is a diagram showing an example of an operation screen operated when creating a layout image. [Figure 18B] 10A and 10B are diagrams showing modified examples of an operation screen operated when creating a layout image. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the configurations described in the following embodiments are merely examples, and the scope of the present invention is not limited to the configurations described in the embodiments. For example, each component constituting the present invention can be replaced with any component that can perform the same function. Furthermore, any component may be added.
[0010] FIG. 1 is a block diagram showing the hardware configuration of a layout image generation system. As shown in FIG. 1, the layout image generation system 100 includes a photographing device (image capture device) 110, an information processing device 120, and a printing device 130, which are communicatively connected to one another via a network 140. The network 140 may be a wide area network (WAN) or a local area network (LAN). Note that the communication connection is not limited to a connection via the network 140, and may be, for example, Bluetooth communication or a USB wired connection. In this embodiment, the photographing device 110 is a digital video camera capable of capturing video. The photographing device 110 is used to capture a person as a subject. This results in acquisition of image data (video data) 900A containing an image of the person. The image data 900A is data composed of multiple frames. Note that the number of people included in the image data 900A may be one or multiple. In addition, when capturing a video, the photographing device 110 also acquires audio data 900B collectively associated with the multiple frames constituting the image data 900A. The audio data 900B mainly contains the human voices of people. The human voices contained in the audio data 900B may be the human voice of one person or may be the human voices of multiple people. Furthermore, the image capturing device 110 also acquires person information (person identification data) 900C that identifies the person contained in the image data 900A. The image data 900A, audio data 900B, and person information 900C are stored in the storage 207 of the image capturing device 110 (see FIG. 2). The image capturing device 110 is configured to be automatically pan- and tilt-controllable, and is used by being placed on, for example, an indoor table or shelf. In this placed state, the image capturing device 110 can automatically detect and track the faces of people present indoors based on the person information 900C. The person information 900C will be described later with reference to FIG. 9. The image capturing device 110 can transmit the image data 900A, audio data 900B, and person information 900C to the information processing device 120 via the network 140. As a result, the information processing device 120 can receive and acquire the image data 900A, the audio data 900B, and the person information 900C.
[0011] The information processing device 120 is, for example, a desktop or notebook personal computer, a tablet terminal, a smartphone, or the like. The information processing device 120 may also function as a cloud server. The information processing device 120 can generate a layout image 1200 (see FIG. 12(a)) or a layout image 1210 (see FIG. 12(b)). The layout image 1200 or the layout image 1210 is an image in which image data 900A and text data in which voice data 900B is converted into text using voice recognition technology are arranged. These images are stored in the USB storage 411 of the information processing device 120 or transmitted to the printing device 130 via the network 140. Upon receiving the layout image 1200 or the layout image 1210, the printing device 130 can print the image. The printing device 130 is a device that forms images, characters, or the like on printing paper or the like using toner or ink. The printing device 130 is not particularly limited, and may be, for example, an MFP or an SMP. When the printing device 130 receives a print job (print instruction) from the information processing device 120, it analyzes the data included in the print job and executes image processing for printing. After this execution, the printing device 130 prints on printing paper or the like.
[0012] FIG. 2 is a block diagram showing the hardware configuration of the photographing device. As shown in FIG. 2, the photographing device 110 has a network I / F (Interface) 201, a CPU 202, a ROM 203, a RAM 204, a camera 205, a microphone 206, a storage 207, and a USB_I / F 208. The network I / F 201 to the USB_I / F 208 are connected via a system bus 210 so as to be able to communicate with each other, i.e., to be able to send and receive data. The network I / F 201 sends and receives various types of data to and from external devices such as the information processing device 120 via the network 140. The CPU 202 is a controller (computer) for controlling the entire photographing device 110. The CPU 202 stores non-volatile memory The OS is started by a startup program stored in ROM 203, which is a memory. The CPU 202 also executes a control program recorded in storage 207 on the OS to control the entire image capturing apparatus 110. This control program includes, for example, a program for causing the CPU 202 to execute each unit and each means of the image capturing apparatus 110. The RAM 204 operates as a temporary storage area such as the main memory or work area of the CPU 202. The storage 207 is a non-volatile memory such as a writable and readable HDD or SSD, and can store, in addition to the control program, for example, image data 900A, audio data 900B, etc.
[0013] The camera 205 can capture moving images and still images. This generates image data 900A. EXIF information, such as the date and time of capture and the location of capture, is attached to the image data 900A. The microphone 206 can convert audio input via the microphone 206 into a digital signal. This generates audio data 900B. EXIF information, such as the date and time of capture and the location of capture, is attached to the audio data 900B. In the case of capturing moving images, the image data 900A and the audio data 900B are associated with each other based on the EXIF information. The USB_I / F 208 is an interface that can be connected to a USB storage 209 via serial communication. This allows the image data 900A, the audio data 900B, and the like to be stored in the USB storage 209 as well. The image capturing device 110 can also be connected to an external device, such as the information processing device 120, via the USB_I / F 208.
[0014] FIG. 3 is a block diagram showing the software configuration of the photographing device. As shown in FIG. 3, the photographing device 110 has a device control unit 301, a network control unit 302, a personal information registration unit 303, a camera control unit 304, a microphone control unit 305, a storage control unit 306, and a USB control unit 307. These software functions are realized by the CPU 202 executing programs loaded into the RAM 204. The device control unit 301 has a function of controlling the network control unit 302 to the USB control unit 307, thereby controlling the operation of the photographing device 110 as a whole. The network control unit 302 has a function of communicating with the network 140 via the network I / F 201 and transmitting and receiving various data to and from external devices. This allows the photographing device 110 to transmit, for example, image data 900A, audio data 900B, and personal information 900C to the information processing device 120. The personal information registration unit 303 has a function of registering personal information 900C, which is characteristic information for identifying a person, based on face image data, audio data 900B, etc. included in the image data 900A. As described above, the image capturing device 110 can automatically detect and track the face of a person based on the person information 900C. The camera control unit 304 has a function of controlling the camera 205 to capture an image with the camera 205. The microphone control unit 305 has a function of controlling the microphone 206 to collect sound with the microphone 206. The storage control unit 306 has a function of controlling the storage 207 to read and write various data from and to the storage 207. The USB control unit 307 has a function of controlling an external device and a USB storage 209 connected via a USB I / F 208.
[0015] 4 is a block diagram showing the hardware configuration of the information processing device. As shown in FIG. 4, the information processing device 120 includes a network I / F 401, a CPU 402, a ROM 403, a RAM 404, a storage 405, an input device I / F 406, a display device I / F 408, and a USB I / F 410. The network I / Fs 401 to 410 are communicably connected to one another via a system bus 420. The network I / F 401 transmits and receives various data to and from external devices such as the photographing device 110 and the printing device 130 via the network 140. The CPU 402 is a controller (computer) for controlling the entire information processing device 120. The CPU 402 starts up the OS by a boot program stored in the ROM 403, which is a non-volatile memory. The CPU 402 can also execute a control program recorded in the storage 405 on the OS to control the entire information processing device 120. This control program includes, for example, a program for causing the CPU 402 to execute each unit and each means (control method of the information processing device) of the information processing device 120. The RAM 404 operates as a temporary storage area such as the main memory or work area of the CPU 402. The storage 405 is a non-volatile memory such as a writable HDD or SDD, and can store, in addition to the control program described above, for example, image data 900A and audio data 900B transmitted from the image capturing device 110.
[0016] The input device I / F 406 is an interface connectable to an input device (operation means) 407. The input device 407 is a device through which a user performs input operations such as operation instructions to the information processing device 120. The input device 407 is not particularly limited and examples thereof include a mouse and a keyboard. The display device I / F 408 is an interface connectable to a display device (display means) 409. The display device 409 is a device that displays various information such as images and text. The display device 409 is not particularly limited and examples thereof include a liquid crystal display. The USB_I / F 410 is an interface connectable to a USB storage 411 via serial communication. This allows image data 900A, audio data 900B, etc. transmitted from the image capturing device 110 to be stored in the USB storage 411 as well. The information processing device 120 can also be connected to external devices such as the image capturing device 110 and the printing device 130 via the USB_I / F 410.
[0017] FIG. 5 is a block diagram showing the software configuration of the information processing device. As shown in FIG. 5, the information processing device 120 includes a device control unit 501, a network control unit (acquisition means) 502, a storage control unit 503, a USB control unit 504, a display control unit 505, and an input control unit 506. The information processing device 120 also includes an image data analysis unit 507, an audio data analysis unit 508, an audio data conversion unit (conversion means) 509, a text analysis unit 510, a layout image generation unit (generation means) 511, and a print instruction unit 512. These software functions are realized by the CPU 402 executing programs loaded into the RAM 404. The device control unit 501 has a function of controlling the network control unit 502 to the print instruction unit 512, thereby controlling the operation of the information processing device 120 as a whole. The network control unit 502 has a function of communicating with the network 140 via the network I / F 401 and transmitting and receiving various data to and from external devices. As a result, the information processing device 120 can receive and acquire image data 900A, audio data 900B, person information 900C, etc. from the imaging device 110 (acquisition process). Furthermore, the information processing device 120 can transmit a print job, etc., including a layout image 1200, etc., to the printing device 130. The storage control unit 503 has a function of controlling the storage 405 to read and write various data from and to the storage 405. The USB control unit 504 has a function of controlling an external device and a USB storage 411 connected via a USB I / F 410. The display control unit 505 has a function of controlling a display device 409 connected via a display device I / F 408. This enables various information to be displayed on the display device 409. The input control unit 506 has a function of controlling an input device 407 connected via an input device I / F 406. This enables operation instructions and the like to be input via the input device 407.
[0018] The image data analysis unit 507 has a function of analyzing the image data 900A. For example, machine learning or image analysis technology is used to analyze the image data 900A. As a result of this analysis, for example, person identification data for identifying a specific person included in the image data 900A is obtained. The audio data analysis unit 508 has a function of analyzing the audio data. This allows multiple speakers (people) to be identified based on the characteristics of their voices and the audio data 900B to be divided into segments for each speaker. For example, deep learning-based speaker identification technology or sound source separation technology can be used to analyze the audio data. The audio data conversion unit 509 has a function of analyzing the content of the audio data 900B and converting the audio data 900B into text data (conversion step). The audio data conversion unit 509 can also convert each of the data divided for each speaker into text data. For example, deep learning-based voice recognition technology can be used to convert the text. The text analysis unit 510 has a function of analyzing the content of the text data and dividing sentences according to context and phrases. For example, natural language processing technology using deep learning can be used to divide the text. The layout image generation unit 511 has a function of generating a layout image (for example, layout image 1200 or layout image 1210) in which specific text data included in the text data and image data 900A are visualized and arranged (generation step). The print instruction unit 512 has a function of issuing a print instruction to the printing device 130 via the network I / F 401. Note that when the printing device 130 is connected to the USB_I / F 410, the print instruction unit 512 can issue a print instruction to the printing device 130 via the USB_I / F 410.
[0019] As described above, the information processing device 120 can use deep learning as a machine learning algorithm, but is not limited to this. For example, a support vector machine, a logistic regression, a decision tree, or the like may also be used. Furthermore, the information processing device 120 may further include a GPU as a hardware configuration. The GPU is a processor for neural network calculations. When the information processing device 120 further includes a GPU, the GPU or the CPU 402 may operate, or the CPU 402 may operate, or both may operate, depending on the type of processing executed by the information processing device 120. Furthermore, the information processing device 120 may be equipped with a TPU or the like instead of a GPU.
[0020] 6 is a block diagram showing the hardware configuration of the printing device. As shown in FIG. 6, the printing device 130 has a network I / F 601, a CPU 602, an eMMC 603, a ROM 604, a RAM 605, a storage 606, a USB I / F 607, a paper feed tray I / F 609, an operation unit 611, an image processing unit 612, and a printer 613. The network I / F 601 to the USB I / F 607, the paper feed tray I / F 609, and the operation unit 611 to the printer 613 are communicably connected to one another via a system bus 620. The network I / F 601 transmits and receives various data to and from external devices such as the information processing device 120 via the network 140. The CPU 602 is a controller (computer) for controlling the entire printing device 130. The CPU 602 starts up the OS using a startup program stored in the ROM 604, which is a non-volatile memory. The CPU 602 also executes a control program recorded in the storage 606 on the OS to control the entire printing device 130. The eMMC 603 is configured as a flash memory and can store the control program for the CPU 602. The RAM 605 operates as a temporary storage area such as the main memory or work area of the CPU 602. The storage 606 is a writable and readable non-volatile memory such as a HDD or SDD, and can store the control program and the like. The USB_I / F 607 is an interface that can be connected to a USB storage 608 via serial communication. This allows print data and the like to be stored in the USB storage 608.
[0021] The paper feed tray I / F 609 is an interface connectable to a paper feed tray 610. The paper feed tray 610 can feed the printing paper required for printing by the printing device 130 one sheet at a time to the printer 613. The operation unit 611 is a device that allows a user to input operation instructions and the like to the printing device 130. The operation unit 611 has an input device such as a keyboard. The operation unit 611 also functions as a display unit that displays various information. In this case, it is preferable that the operation unit 611 has a display with a touch panel function. The image processing unit 612 is a hardware module that performs image processing such as decoding of print data and enlargement / reduction. The printer 613 is a device that prints using toner or ink on printing paper supplied from the paper feed tray 610.
[0022] FIG. 7 is a block diagram showing the software configuration of a printing device. As shown in FIG. 7, the printing device 130 has a device control unit 701, a network control unit 702, a storage control unit 703, a USB control unit 704, an operation control unit 705, a paper feed tray control unit 706, and a printer control unit 707. These software functions are realized by the CPU 602 executing programs loaded into the RAM 605. The device control unit 701 has a function of controlling the network control unit 702 to the printer control unit 707, thereby controlling the operation of the printing device 130 as a whole. The network control unit 702 has a function of communicating with the network 140 via the network I / F 601 and transmitting and receiving various data to and from external devices. This allows the printing device 130 to receive print jobs and the like from the information processing device 120. The storage control unit 703 has a function of controlling the storage 606 and reading and writing various data from and to the storage 606. The USB control unit 704 has a function of controlling the external device and the USB storage 608 connected via the USB I / F 607. The operation control unit 705 has a function of controlling the operation unit 611 and acquiring input information input from the operation unit 611. The paper feed tray control unit 706 has a function of controlling the paper feed tray 610 via the paper feed tray I / F 609 and supplying printing paper from the paper feed tray 610. The printer control unit 707 has a function of controlling the printer 613 and causing the printer 613 to print.
[0023] Fig. 8 is a flowchart showing the processing executed by the image capturing device. As shown in Fig. 8, in step S800, the CPU 202 of the image capturing device 110 registers, in storage 207, person information 900C of the person to be captured by the image capturing device 110, i.e., the person who will be the subject. The person information 900C is person identification data that identifies the person. The person identification data is not particularly limited, and may include, for example, at least one of image data of the person's face and audio data of the person's actual voice.
[0024] In step S801, CPU 202 controls camera 205 to capture an image of a specific person identified based on person information 900C. As a result, image data 900A is obtained and stored in storage 207. As described above, image capturing device 110 can automatically detect and track the face of a specific person. This allows image capturing device 110 to focus on the face of the specific person and continue to automatically capture images of the specific person. Note that image capturing device 110 can also continue to capture images of the specific person through user operation.
[0025] In step S802, if CPU 202 determines that the specific person has uttered their real voice, it records audio data 900B containing the specific person's real voice. As a result, audio data 900B is stored in storage 207. Note that in image capture device 110, recording of audio data 900B can also be started by a user operation. Also, steps S801 and S802 may be executed in reverse order or simultaneously.
[0026] In step S803, the CPU 202 transmits the image data 900A and the audio data 900B stored in the storage 207 to the information processing device 120 via the network I / F 201. As a result, the information processing device 120 acquires the image data 900A and the audio data 900B.
[0027] In step S804, the CPU 202 transmits all of the personal information 900C stored in the storage 207 to the information processing device 120 via the network I / F 201. As a result, the personal information 900C is acquired in the information processing device 120. Note that steps S803 and S804 may be executed in reverse order or may be executed simultaneously.
[0028] FIG. 9 is a diagram illustrating an example of person information stored in the imaging device. Person information 900C illustrated in FIG. 9 is information related to identifying a person included in image data 900A. In this embodiment, this information includes a person ID 901, image feature information 902, and audio feature information 903. The person ID 901 is a code for identifying a person included in image data 900A. While the alphabet is used as the person ID 901 in this embodiment, the person ID is not limited to this and may be, for example, letters, symbols, or a combination thereof. The image feature information 902 is a feature of the image of the person included in image data 900A. Although the image feature information 902 in this embodiment uses a face image of the person, the image feature information is not limited to this and may be, for example, the person's physique. In addition, although the file format of the image feature information 902 is "JPEG" in this embodiment, the file format is not limited to this and may be, for example, "PNG." The image feature information 902 may include multiple features for each person. In this case, a file including multiple features may be linked to the image feature information 902. The audio feature information 903 is a feature of the voice of a person included in the audio data 900B related to the image data 900A. In this embodiment, the audio feature information 903 uses the person's real voice data, but is not limited to this. In this embodiment, the file format of the audio feature information 903 is "MP3," but is not limited to this and may be, for example, "WAV." In addition, the audio feature information 903 may include multiple features per person. In this case, a file including multiple features may be linked to the audio feature information 903.
[0029] FIG. 10A is a flowchart showing processing executed by the information processing device. FIG. 10B is a conceptual diagram illustrating an example of processing executed by the information processing device. As shown in FIG. 10A, in step S1000, the CPU 402 of the information processing device 120 receives (acquires) various data from the image capturing device 110 via the network control unit 502. The various data include the image data 900A and audio data 900B (see FIG. 10B) transmitted in step S803 and the person information 900C transmitted in step S804. The image data 900A, audio data 900B, and person information 900C are saved in the storage 405 of the information processing device 120. In step S1000, for example, by operating the operation screen 1800 (see FIG. 18) to specify the image capturing device 110 and the shooting date from which data is to be acquired, the image data 900A and audio data 900B that satisfy the specified conditions can be acquired.
[0030] In step S1001, CPU 402 converts voice data 900B stored in storage 405 in step S1000 into text information (text data) 1400 (see FIG. 14) using voice data analysis unit 508 and voice data conversion unit 509. FIG. 14 is a diagram showing an example of text information. As shown in FIG. 14, text information 1400 includes person ID 1401, start time 1402, end time 1403, and divided text 1404. Person ID 1401 to divided text 1404 will be described later with reference to FIG. 14. Also, detailed processing (text conversion processing) executed in step S1001 will be described later with reference to FIG. 13.
[0031] In step S1002, CPU 402 selects one image data 900A1 (frame) to be included in layout image 1010 from the image data 900A (multiple frames) saved in storage 405 in step S1000 (see FIG. 10B). This selection is performed, for example, by displaying operation screen 1810 on display device 409, on which image data 900A can be confirmed, and allowing the user to select desired image data 900A1 on operation screen 1810 via input device 407 (selection step). Thus, in this embodiment, operation screen 1810 (input device 407) functions as a selection means for selecting image data 900A. Note that the number of image data 900A1 selected on operation screen 1810 can be one or more.
[0032] In step S1003, CPU 402 causes image data analysis unit 507 to extract a facial image (face region) of a person included in image data 900A1 selected in step S1002. This extraction is performed by extracting the face of the person included in image data 900A1 based on person information 900C stored in storage 405 in step S1000. Then, specific text information 1020, which is a text version of the extracted person's voice, is further extracted from text information 1400 acquired in step S1001. As described above, in this embodiment, image data analysis unit 507 functions as extraction means for extracting specific text information 1020. In particular, when image data 900A is video data, the same person included in image data 900A1 selected in step S1002 and image data 900A2 and 900A3 following image data 900A1 is extracted (see FIG. 10B). It is then preferable to extract a series of text information of the person's voice from the text information 1400 as the specific text information 1020. This allows for more accurate extraction of the specific text information 1020 in the case of video data. Note that image data 900A2 and image data 900A3 are image data that follow image data 900A1, but depending on the image data 900A1 selected in step S1002, they may be either before or after. The detailed processing (face extraction processing) performed in step S1003 will be described later with reference to FIG. 15.
[0033] In step S1004, CPU 402 generates layout image 1010 using layout image generation unit 511. Layout image 1010 is an image in which specific text information 1020 extracted in step S1003 and image data 900A1 selected in step S1002 are combined and arranged (see FIG. 10B). Details of the process (layout image generation process) executed in step S1004 will be described later with reference to FIG. 17.
[0034] In step S1005, CPU 402 determines whether layout image generation processing has been completed for all image data 900A selected in step S1002. If it is determined in step S1005 that layout image generation processing has been completed, processing proceeds to step S1006. On the other hand, if it is determined in step S1005 that layout image generation processing has not been completed, processing returns to step S1003, and subsequent steps are executed in order.
[0035] In step S1006, CPU 402 determines whether or not a command to print layout image 1010 generated in step S1004 has been issued. This determination is made, for example, based on whether or not print button 1826 has been selected on preview screen 1820 (see FIG. 18A). If print button 1826 has been selected, a print command is issued; if print button 1826 has not been selected, a print command is not issued. If the determination in step S1006 determines that a print command has been issued, the process proceeds to step S1007. On the other hand, if the determination in step S1006 determines that a print command has not been issued, the process proceeds to step S1008.
[0036] In step S1007, the CPU 402 transmits a print instruction for the layout image 1010 to the printing device 130. As a result, the printing device 130 prints the layout image 1010.
[0037] In step S 1008 , the CPU 402 stores the layout image 1010 in the storage 405 .
[0038] 11 is a flowchart showing the processing executed by the printing device 130. As shown in FIG. 11, in step S1100, the CPU 602 of the printing device 130 controls the network control unit 702 to receive a print instruction from the information processing device 120.
[0039] In step S1101, the CPU 602 selects the paper feed tray 610 based on the print instruction received in step S1100. The paper feed tray 610 contains printing paper of the paper size specified in the print instruction.
[0040] In step S1102, the CPU 602 performs image processing on the image data included in the print instruction received in step S1100 to convert it into binary image data.
[0041] In step S1103, the CPU 602 transports printing paper from the paper feed tray 610 selected in step S1101 to the printer 613, and causes the printer 613 to print the image data that was image processed in step S1102. This results in a printed matter on which, for example, the layout image 1200 (see FIG. 12(a)) or the layout image 1210 (see FIG. 12(b)) is printed.
[0042] FIG. 12 is a diagram showing an example of a layout image. FIG. 12(a) is a first layout image as an example of a layout image. FIG. 12(b) is a second layout image as an example of a layout image. The layout image generation unit 511 is capable of generating a layout image 1200 shown in FIG. 12(a) and a layout image 1210 shown in FIG. 12(b). As shown in FIG. 12(a), in the layout image 1200, an image of specific text information is arranged as a speech bubble on an image of a person. In this embodiment, the layout image 1200 includes a person 1201, a person 1202, specific text information 1203, and specific text information 1204. Near the face of the person 1201, specific text information 1203 extracted by the image data analysis unit 507 is attached as a speech bubble of the real voice spoken by the person 1201. Near the face of the person 1202, specific text information 1204 extracted by the image data analysis unit 507 is attached as a speech bubble of the person's voice.
[0043] As shown in FIG. 12(b), in layout image 1210, images of specific text information are arranged in adjacent columns vertically or horizontally to an image of a person. In this embodiment, layout image 1210 includes image data 1211 and specific text information 1212. The specific text information 1212 is a text, such as a bulleted list, of the voice of a person 1213 included in the image data 1211, and is arranged adjacent to the image data 1211 below. In the configuration shown in FIG. 12(b), the specific text information 1212 is arranged below the image data 1211, but is not limited to this and may be arranged above, to the left, or to the right of the image data 1211, for example. Note that the layout image is not limited to layout image 1200 and layout image 1210 and may be, for example, an image combining layout image 1200 and layout image 1210.
[0044] FIG. 13 is a flowchart showing detailed processing (text conversion processing) in step S1001, which is a subroutine of the flowchart shown in FIG. 10A. As shown in FIG. 13, in step S1300, the CPU 402 of the information processing device 120 causes the voice data analysis unit 508 to separate the real voices of all people, i.e., all speakers, included in the voice data 900B into individual speakers. Specifically, the voice data analysis unit 508 identifies each speaker based on the characteristics of each speaker's voice using, for example, deep learning-based speaker identification technology or sound source separation technology, and divides the voice data 900B into individual speakers. This makes it possible to extract the voice data 900B of the desired speaker's voice when the voices of multiple speakers are included in the voice data 900B, thereby improving the accuracy of processing from step S1300 onward.
[0045] In step S1301, the CPU 402 identifies the speaker included in the voice data 900B separated in step S1300 using the voice data analysis unit 508. Specifically, the voice data analysis unit 508 identifies the speaker included in the voice data 900B based on the person ID 901 and voice feature information 903 of the person information 900C, for example, using a speaker identification technology using deep learning.
[0046] In step S1302, CPU 402 converts the voice data 900B for each speaker identified in step S1301 into text using voice data conversion unit 509. Specifically, voice data conversion unit 509 converts the voice data 900B for each speaker into text information 1400 using, for example, deep learning-based voice recognition technology. The text information 1400 includes timestamp information at regular intervals from the start time of the EXIF information of the voice data 900B. This makes it possible to analyze when words in the text information 1400 were spoken. After this analysis, a process is performed to associate the person ID 901 associated with the voice data 900B with the converted text information 1400.
[0047] In step S1303, the CPU 402 causes the text analysis unit 510 to divide the text information 1400 for each speaker acquired in step S1302 into sentences. Specifically, the text analysis unit 510 analyzes the content of the text information 1400 for each speaker acquired in step S1302 using, for example, a natural language processing technique based on deep learning. The text analysis unit 510 then divides the text information 1400 into divided texts 1404 in sentence units according to the context. Also, in step S1303, a process is executed to associate a start time 1402 and an end time 1403 of an utterance with the divided texts 1404 based on the timestamp information of the text information 1400.
[0048] In step S1304, CPU 402 determines whether or not processing has been completed for all audio data 900B. If the determination in step S1304 indicates that processing has been completed for all audio data 900B, processing ends. On the other hand, if the determination in step S1304 indicates that processing has not been completed for all audio data 900B, processing returns to step S1300, and subsequent steps are executed in order.
[0049] FIG. 14 is a diagram showing an example of text information. Text information 1400 shown in FIG. 14 is acquired in step S1001 of the flowchart in FIG. 10A. The text information 1400 includes a person ID 1401, a start time 1402, an end time 1403, and segmented text 1404, which are associated with one another. The person ID 1401 is the person ID 901 in the person information 900C. The start time 1402 is the time when the person assigned with each person ID 901 started to speak. The end time 1403 is the time when the person assigned with each person ID 901 finished speaking. The start time 1402 and the end time 1403 are generated based on the EXIF information of the audio data 900B received in step S1000. The segmented text 1404 is text information segmented into sentences in step S1303. The format of the divided text 1404 is not particularly limited, and may be, for example, a character string data format or a file format saved with an extension such as ".txt."
[0050] FIG. 15 is a flowchart showing detailed processing (face extraction processing) in step S1003, which is a subroutine of the flowchart shown in FIG. 10A. As shown in FIG. 15, in step S1500, the CPU 402 of the information processing device 120 extracts a face region of the image of a person included in image data 900A using the image data analysis unit 507, and acquires coordinate information 1604 of the face region (see FIG. 15). Specifically, the image data analysis unit 507 extracts the face region of the image of a person included in image data 900A as a rectangular region using, for example, machine learning or image analysis technology. Then, the image data analysis unit 507 acquires the coordinates of two diagonal points of the rectangular region as coordinate information 1604 of the face region. Note that the face region of the image of a person is not limited to being acquired using the coordinate information 1604.
[0051] In step S1501, the CPU 402 identifies the person having the coordinate information 1604 acquired in step S1500 using the image data analysis unit 507. Specifically, the image data analysis unit 507 determines whether the identified person matches the person having the image feature information 902 linked to the person ID 901, for example, using an image analysis technique in machine learning.
[0052] In step S1502, the CPU 402 associates the coordinate information 1604 and the like with the person identified in step S1501.
[0053] FIG. 16 is a diagram showing an example of face area information. As shown in FIG. 16(a), face area information 1600 includes an image data name 1601, a shooting time 1602, a person ID 1603, and coordinate information 1604. The image data name 1601 is the file name of the image data 900A1 selected in step S1002 (see FIG. 10A). The shooting time 1602 is the time when the image data 900A1 to which the image data name 1601 is assigned was shot. The shooting time 1602 is acquired based on the EXIF information of the image data 900A1. The person ID 1603 is a code for identifying the person identified in step S1501 (see FIG. 15). For example, as described above, if the person to be identified matches the person having the image feature information 902 linked to the person ID 901, the person ID 901 can be used as the person ID 1603. On the other hand, if the identified target person does not match the person having the image feature information 902 linked to the person ID 901, "NULL" can be set as the person ID 1603. The coordinate information 1604 is the coordinate information of the face area acquired in step S1500 (see FIG. 15). FIG. 16(b) shows image data 1605 having an image data name 1601 of "sample1.jpg". In the image data 1605, the pair of (50,50) and (100,100) and the pair of (150,150) and (200,200) are depicted as the coordinate information 1604. FIG. 16(c) shows image data 1606 having an image data name 1601 of "sample2.jpg". In the image data 1606, the pair of (200,50) and (250,100) are depicted as the coordinate information 1604.
[0054] Fig. 17 is a flowchart showing detailed processing (layout image generation processing) in step S1004, which is a subroutine of the flowchart shown in Fig. 10A. As shown in Fig. 17, in step S1700, CPU 402 of information processing device 120 acquires, from face region information 1600, shooting time 1602 of image data 900A to be used for generating a layout image, i.e., image data 900A to be included in the layout image.
[0055] In step S1701, CPU 402 extracts text information recorded within a certain time before and after the shooting time 1602 acquired in step S1700. Specifically, CPU 402 references start time 1402 and end time 1403 of text information 1400, and extracts all text information 1400 recorded within a certain time (for example, one minute) before and after the shooting time 1602. Note that it is preferable that the certain time before and after the shooting time 1602, which is the middle of the time, be changeable as appropriate.
[0056] In step S1702, the CPU 402 determines whether generation of a first layout image (see FIG. 12A), i.e., a layout image including speech bubbles, is valid as a layout image. As described above, in this embodiment, the CPU 402 functions as a determination unit that determines whether generation of a first layout image is valid. Note that in the information processing device 120, a section that functions as a determination unit may be provided separately from the CPU 402. The determination in step S1702 is made, for example, based on whether a check box 1831 on the operation screen 1830 (see FIG. 18A) is checked. If the check box 1831 is checked, it is determined that generation of a first layout image is valid. If the check box 1831 is not checked, it is determined that generation of a first layout image is not valid. Then, if it is determined that generation of a first layout image is valid as a result of the determination in step S1702, the process proceeds to step S1703. On the other hand, if it is determined in step S1702 that generation of the first layout image is not valid, the process proceeds to step S1706.
[0057] In step S1703, CPU 402 determines whether the speaker of the human voice included in the text information extracted in step S1701 is included in image data 900A, i.e., whether the speaker is shown in the image data. Specifically, CPU 402 determines whether any data among person IDs 1401 in text information 1400 extracted in step S1701 matches person ID 1603 shown in image data 900A to be processed. If the determination in step S1703 indicates that the speaker is included in image data 900A, the process proceeds to step S1704. On the other hand, if the determination in step S1703 indicates that the speaker is not included in image data 900A, the process proceeds to step S1706.
[0058] In step S1704, the CPU 402 selects text information to be included in the first layout image. Specifically, the CPU 402 selects segmented text 1404 based on a selection algorithm from among the text information in which the person ID 1603 appearing in the image data to be processed matches the person ID 1401 in the text information 1400 extracted in step S1701. The "selection algorithm" refers to, for example, a method of selecting one piece of text in which the conversation start time 1402 is closest to the shooting time 1602, or a method of selecting on an operation screen (not shown) displaying candidate segmented text 1404 in step S1704. Note that if there are multiple person IDs 901 that match the face area information 1600 in the text information 1400 extracted in step S1701, one or more segmented texts 1404 can be selected for each person ID 1603. This allows each person's remarks to be displayed as speech bubbles on the first layout image.
[0059] In step S1705, CPU 402 places divided text 1404 selected in step S1704 near the face area of the person on image data 900A, and causes layout image generation unit 511 to generate a first layout image. Specifically, layout image generation unit 511 creates objects that surround divided text 1404 with speech bubbles. Thereafter, layout image generation unit 511 places the objects side by side near the face area of the person with person ID 1603 that is the same as person ID 1401 of the speaker. In this way, a first layout image is generated.
[0060] In step S1706, CPU 402 generates a second layout image (see FIG. 12(b)) as a layout image using layout image generation unit 511. This second layout image is an image in which text information 1400 (divided text 1404) extracted in step S1701 and image data 900A are arranged either vertically or horizontally. For example, all of the text information 1400 arranged in order of earliest start time 1402 may be arranged, or divided text 1404 with the same person ID 1401 may be extracted and arranged in separate locations for each speaker.
[0061] As described above, before generating a layout image, the information processing device 120 determines whether generation of either the first layout image or the second layout image is valid as the layout image. If it is determined that generation of the first layout image is valid as a result of this determination, the first layout image is generated, and if it is determined that generation of the second layout image is valid, the second layout image is generated. In this way, when desired image data 900A to be included in the layout image is selected, the information processing device 120 can acquire a first layout image or a second layout image including the image data 900A and the text information 1400. As a result, when the layout image generation system 100 is used in, for example, a nursery school or an elementary school, memorable image data can be selected as the image data 900. Then, the first layout image (or the second layout image) is an image including an image of a kindergartener or elementary school student included in the memorable image data and text information of the child's voice. This first layout image can be published as a memorable image in, for example, a class newsletter or a yearbook.
[0062] FIG. 18A is a diagram showing an example of an operation screen operated when creating a layout image. The operation screen 1800 shown in FIG. 18A(a) is a data reading screen operated when the information processing device 120 receives various data from the image capturing device 110, and is displayed on the display device 409 under the control of the display control unit 505. The operation screen 1800 includes a device selection section 1801, a shooting date selection section 1802, and a start button 1803. The device selection section 1801 allows the user to select the name of the image capturing device 110 connected to the information processing device 120 as the source of image data 900A and the like. The shooting date selection section 1802 allows the user to set a shooting date period and specify image data 900A acquired during that period. By operating, i.e., pressing, the start button 1803, the image data 900A and audio data 900B of the shooting date specified in the shooting date selection section 1802 can be received from the image capturing device 110 selected in the device selection section 1801. After receiving this, the screen transitions to an operation screen 1810 shown in FIG. 18A(b).
[0063] The operation screen 1810 is an image data selection screen that allows the user to select desired image data to be included in the layout image from the image data 900A. The operation screen 1810 includes a list display area 1811, a set button 1812, and a confirm button 1813. All image data 900A received by operating the start button 1803 on the operation screen 1800 is displayed in the list display area 1811. The user can select desired image data to be included in the layout image from all image data 900A, for example, by clicking with a mouse. After this selection, the user operates the set button 1812 to transition to an operation screen 1830 shown in FIG. 18A(d). After the desired image data is selected, the user operates the confirm button 1813 to start a layout image generation process for the selected image data. After the layout image generation process is completed, the screen transitions to a preview screen 1820 shown in FIG. 18A(c).
[0064] The preview screen 1820 includes a preview area 1821, a previous image button 1822, a next image button 1823, an edit button 1824, a save button 1825, and a print button 1826. The preview area 1821 displays a preview image of the layout image acquired in the layout image generation process. This allows the user to check what the layout image will look like. If there are multiple layout images, the multiple layout images can be displayed in sequence in the preview area 1821 by operating the previous image button 1822 or the next image button 1823. The layout images can be saved in the storage 405 by operating the save button 1825. If there are multiple layout images, the user may be able to select whether to save only the layout image displayed as a preview image in the preview area 1821 or all of the layout images. The saving of the layout images is not limited to being performed by operating the save button 1825. For example, the saving may be performed automatically after a predetermined time has elapsed since the preview screen 1820 was displayed, regardless of whether the save button 1825 was operated. By operating the print button 1826, it is possible to instruct the printing device 130 to print the layout image. As a result, printing of the layout image is executed in the printing device 130. Note that if there are multiple layout images, it may be possible to select whether to print only the layout image displayed as a preview image in the preview area 1821, or to print all of the layout images.
[0065] 18A(d) includes a check box 1831 and a save button 1832. As described above, if the check box 1831 is checked, it is determined that generation of the first layout image is valid, and if the check box 1831 is not checked, it is determined that generation of the first layout image is not valid. By operating the save button 1832, information regarding whether the check box 1831 is checked or not is saved in the storage 405.
[0066] FIG. 18B illustrates a modified example of an operation screen used when creating a layout image. A preview area 1821 of a preview screen 1820 shown in FIG. 18B(a) displays a preview image 1840 of a layout image acquired in the layout image generation process. The preview image 1840 is an image that mimics the first layout image in which an image 1841 of image data 900A and speech bubbles 1842 and 1843, which are images of specific text information 1020, are arranged. Speech bubble 1842 contains the text "aaa," and speech bubble 1843 contains the text "bbb." Suppose you want to change the text content of speech bubble 1843 from "bbb" to "ccc." Therefore, by operating edit button 1824, speech bubble 1843 becomes editable. In this editable state, you can input "ccc" using input device 407. As a result, the preview screen 1820 becomes the state shown in FIG. 18B(b). 18B(c), the layout image generating unit 511 can reflect the result of the operation on the input device 407, i.e., the result of inputting "ccc" through the input device 407, in the layout image 1850. Furthermore, the printing device 130 can obtain a printed matter in which the layout image 1850 is printed. Editing in the preview area 1821 (preview image) is not limited to changing the content of the text included in the speech bubble 1842 or the speech bubble 1843. For example, at least one of the following edits (operations) can be performed: changing the positional relationship between the image 1841 and the speech bubbles 1842 and 1843; and deleting the speech bubbles 1842 and 1843.
[0067] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments, and various modifications and variations are possible within the scope of the gist of the present invention. The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or storage medium, and having one or more processors in the computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., an ASIC) that realizes one or more functions.
[0068] The disclosure of this embodiment includes the following configuration, method, and program. (Configuration 1) An acquisition means for acquiring image data including an image of a person and audio data associated with the image data; a conversion means for converting the voice data acquired by the acquisition means into text data; a selection means for selecting the image data acquired by the acquisition means; An information processing device characterized by comprising: a generation means for generating a layout image in which specific text data of the voice uttered by the person contained in the image data selected by the selection means and the image data selected by the selection means are arranged. (Configuration 2) The acquisition means is capable of acquiring person identification data that identifies the person, 2. The information processing apparatus according to configuration 1, further comprising an extraction means for extracting the specific text data from the text data based on the person identification data acquired by the acquisition means. (Configuration 3) The information processing device described in Configuration 2, wherein the extraction means extracts the face of the person included in the image data selected by the selection means from the text data based on the person identification data acquired by the acquisition means, and extracts the specific text data for the person. (Configuration 4) The information processing device according to configuration 2 or 3, characterized in that the generating means generates an image in which the specific text data extracted by the extracting means and the image data selected by the selecting means are arranged as the layout image. (Configuration 5) The image data is video data composed of a plurality of frames, 5. The information processing device according to any one of configurations 1 to 4, wherein the selection means is capable of selecting one frame from the plurality of frames. (Configuration 6) The audio data is audio data collectively associated with the plurality of frames, The information processing device according to configuration 5, further comprising an extraction means for extracting, from the text data, text data of the voice uttered by the person contained in the one frame selected by the selection means and at least one of the frames before and after the frame, as the specific text data. (Configuration 7) The image data includes a plurality of images of the person, 7. The information processing device according to any one of configurations 1 to 6, further comprising extraction means for extracting the specific text data for each of the people. (Configuration 8) The information processing apparatus according to configuration 7, wherein the generating means generates the layout image in which the specified text data is arranged. (Configuration 9) The information processing device described in any one of configurations 1 to 8, characterized in that the generation means is capable of generating, as the layout image, a first layout image in which an image of the specific text data is arranged as a speech bubble on the image of the person, and a second layout image in which an image of the specific text data is arranged as adjacent columns above, below, or to the left and right on the image of the person. (Configuration 10) A determination unit is provided for determining whether generation of either the first layout image or the second layout image is effective as the layout image before the generation unit generates the layout image, The information processing device described in configuration 9 is characterized in that, when the judgment by the judgment means determines that generation of the first layout image is valid, the generation means generates the first layout image, and when the judgment by the judgment means determines that generation of the second layout image is valid, the generation means generates the second layout image. (Configuration 11) The image data includes a plurality of images of the person, The audio data includes data of the actual voice of each person, 11. The information processing device according to any one of configurations 1 to 10, wherein the conversion means separates the audio data into data of the real voice of each person and converts each of the data into the text data. (Configuration 12) The information processing device is communicably connected to an imaging device capable of storing the image data, the audio data, and the person identification data, 3. The information processing apparatus according to configuration 2, wherein the acquisition means acquires the image data, the audio data, and the person identification data from the imaging device. (Configuration 13) A display means capable of displaying a preview image of the layout image; 13. The information processing device according to any one of configurations 1 to 12, further comprising an operation means capable of performing at least one of the following operations on the preview image: changing the positional relationship between the image of the image data and the image of the specific text data; deleting the image of the specific text data; and changing the content of the text included in the image of the specific text data. (Configuration 14) The information processing apparatus according to configuration 13, wherein the generating means reflects the operation result of the operating means in the layout image. (Method 1) A method for controlling an information processing device, comprising: an acquisition step of acquiring image data including an image of a person and audio data associated with the image data; a conversion step of converting the voice data acquired in the acquisition step into text data; a selection step of selecting the image data acquired in the acquisition step; A control method for an information processing device, characterized by comprising a generation process for generating a layout image in which specific text data of the voice uttered by the person contained in the image data selected in the selection process and the image data selected in the selection process are arranged. (Program 1) A program for causing a computer to execute each means of the information processing device according to any one of configurations 1 to 14. [Explanation of symbols]
[0069] 110 Imaging equipment 120 Information processing equipment 502 Network control section 509 Audio Data Conversion Unit 511 Layout image generation unit 900A Image Data 900B audio data 1400 Text Information
Claims
1. an acquisition means for acquiring image data including an image of a person and audio data associated with the image data; a conversion means for converting the voice data acquired by the acquisition means into text data; a selection means for selecting the image data acquired by the acquisition means; An information processing device characterized by comprising: a generation means for generating a layout image in which specific text data of the voice uttered by the person contained in the image data selected by the selection means and the image data selected by the selection means are arranged.
2. the acquisition means is capable of acquiring person identification data that identifies the person, 2. The information processing apparatus according to claim 1, further comprising: an extracting unit that extracts the specific text data from the text data based on the person specifying data acquired by the acquiring unit.
3. The information processing device according to claim 2, characterized in that the extraction means extracts the face of the person included in the image data selected by the selection means from the text data based on the person identification data acquired by the acquisition means, and extracts the specific text data for the person.
4. 3. The information processing apparatus according to claim 2, wherein the generating means generates an image in which the specific text data extracted by the extracting means and the image data selected by the selecting means are arranged as the layout image.
5. The image data is video data composed of a plurality of frames, 2. The information processing apparatus according to claim 1, wherein the selection means is capable of selecting one frame from the plurality of frames.
6. the audio data is audio data collectively associated with the plurality of frames, The information processing device according to claim 5, further comprising an extraction means for extracting, from the text data, text data of the voice uttered by the person contained in the one frame selected by the selection means and at least one of the frames before and after the frame, as the specific text data.
7. the image data includes a plurality of images of the person; 2. The information processing apparatus according to claim 1, further comprising: extraction means for extracting the specific text data for each of the persons.
8. 8. The information processing apparatus according to claim 7, wherein said generating means generates the layout image in which each of said specific text data is arranged.
9. The information processing device according to claim 1, characterized in that the generation means is capable of generating, as the layout image, a first layout image in which an image of the specific text data is arranged as a speech bubble on the image of the person, and a second layout image in which an image of the specific text data is arranged as adjacent columns above, below, or to the left and right on the image of the person.
10. a determining means for determining, prior to the generation of the layout image by the generating means, whether generation of either the first layout image or the second layout image is effective as the layout image, 10. The information processing apparatus according to claim 9, wherein the generating means generates the first layout image when it is determined that generation of the first layout image is valid as a result of the judgment by the judging means, and generates the second layout image when it is determined that generation of the second layout image is valid as a result of the judgment by the judging means.
11. the image data includes a plurality of images of the person; The audio data includes data of the actual voice of each person, 2. The information processing apparatus according to claim 1, wherein said converting means separates said voice data into data of the real voice of each person and converts each of said data into said text data.
12. the information processing device is communicably connected to an imaging device capable of storing the image data, the audio data, and the person identification data; 3. The information processing apparatus according to claim 2, wherein the acquisition means acquires the image data, the audio data, and the person identification data from the imaging device.
13. a display means capable of displaying a preview image of the layout image; 2. The information processing device according to claim 1, further comprising an operation means for performing at least one of the following operations on the preview image: changing the positional relationship between the image of the image data and the image of the specific text data; deleting the image of the specific text data; and changing the content of the text included in the image of the specific text data.
14. 14. The information processing apparatus according to claim 13, wherein the generating means reflects the result of the operation by the operating means in the layout image.
15. A method for controlling an information processing device, comprising: an acquisition step of acquiring image data including an image of a person and audio data associated with the image data; a conversion step of converting the voice data acquired in the acquisition step into text data; a selection step of selecting the image data acquired in the acquisition step; A control method for an information processing device, characterized by comprising a generation process for generating a layout image in which specific text data of the voice uttered by the person contained in the image data selected in the selection process and the image data selected in the selection process are arranged.
16. 2. A program for causing a computer to execute each means of the information processing apparatus according to claim 1.
Citation Information
Patent Citations
Minute creating apparatus, minute creating method, and program
JP2019149083A