Data generation system

The data generation system addresses the lack of engagement in digitizing analog photos by converting text into voice data, offering an entertaining digital archiving solution through voice playback.

JP7711473B2Active Publication Date: 2025-07-23DAI NIPPON PRINTING CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2021126786
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-08-02
Publication Date
2025-07-23
Estimated Expiration
2041-08-02

AI Technical Summary

Technical Problem

Conventional services for digitizing analog photos lack engagement and fail to convert text into voice data, resulting in a lack of interactive and entertaining digital archiving solutions.

Method used

A data generation system that includes an image acquisition unit, text extraction unit, voice data generation unit, and storage unit to digitize and convert text in analog photos into voice data, enabling digital archiving with voice playback.

Benefits of technology

Analog photos are digitized with text converted into voice data, providing an entertaining and interactive digital archiving experience by allowing voice playback alongside image display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007711473000001
    Figure 0007711473000001
  • Figure 0007711473000002
    Figure 0007711473000002
  • Figure 0007711473000003
    Figure 0007711473000003
Patent Text Reader

Abstract

To provide a data generation system which converts an analog photograph into digital data and converts a text in the analog photograph into sound data to thereby digitally archive the data as an image with sound.SOLUTION: A data generation system comprises: an image acquisition unit 11 which acquires a photographed image obtained by capturing an analog photograph P from a user terminal 2; a text extraction unit 12 which extracts a text from the photographed image; a sound data generation unit 13 which converts the text into sound to generate sound data; and a storage unit 10 which stores the photographed image and the sound data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data generation system.

Background Art

[0002] Services are provided to scan and digitize analog photos (photo albums) taken in the past so that they can be stored long-term in the form of electronic data. Conventional services simply digitize analog photos and are not interesting.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present invention is to provide a data generation system that digitizes analog photos, converts text in the analog photos into voice data, and digitally archives them as images with voice.

Means for Solving the Problems

[0005] The data generation system of the present invention includes an image acquisition unit that acquires a photographed image obtained by photographing an analog photo from a user terminal, a text extraction unit that extracts text from the photographed image, a voice data generation unit that converts the text into voice to generate voice data, and a storage unit that stores the photographed image and the voice data.

Effects of the Invention

[0006] According to the present invention, an analog photograph can be digitized, and the text in the analog photograph can be converted into voice data and digitally archived as an image with voice. Since the user terminal can play the voice while displaying the image, a highly entertaining service can be provided to the user.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0008] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0009] As shown in FIG. 1, the data generation system according to the present embodiment includes a server device 1 and a user terminal 2 such as a smartphone, and provides a digital archive service for analog photographs. A user registered for the digital archive service activates the digital camera mounted on the user terminal 2 and takes an analog photograph P with the camera. The user terminal 2 transmits the photographed image to the server device 1. The analog photograph P is, for example, a photo sticker produced by a print sticker machine. Usually, a print sticker machine has not only a photographing function but also a drawing function, and the photo sticker includes text such as handwritten characters. The analog photograph P may be a film photograph with characters handwritten with a pen or the like.

[0010] When the server device 1 receives a captured image from the user terminal 2, it extracts text from the captured image and converts the extracted text into voice data. The server device 1 associates the captured image with the voice data and saves it as an image with voice. This image with voice can be displayed and played on the user terminal 2.

[0011] For example, when the user terminal 2 captures an analog photo P as shown in Fig. 2a, the server device 1 extracts the text "Always a friend" from the captured image, converts the text into voice data, and generates an image with voice. When this image with voice is displayed and played on the user terminal 2, as shown in Fig. 2b, the captured image 21 is displayed on the display of the user terminal 2, and when the user presses the play button 22, the voice "Always a friend" is output. The user can digitally archive the analog photo and enjoy the voice at the same time.

[0012] As described above, according to this embodiment, a digital archive service for analog photos with high entertainment value can be provided to the user.

[0013] Next, the configuration of the server device 1 will be described. The server device 1 is a computer equipped with a communication unit, a storage unit, a CPU, etc. When the CPU executes a program, as shown in Fig. 3, the functions of the image acquisition unit 11, the text extraction unit 12, the voice data generation unit 13, and the playback unit 14 are realized.

[0014] The image acquisition unit 11 acquires a captured image of the analog photo P captured from the user terminal 2.

[0015] The text extraction unit 12 extracts text from the captured image acquired by the image acquisition unit 11. For text extraction, known methods such as OCR (Optical Character Recognition) can be used.

[0016] The voice data generation unit 13 converts the text extracted by the text extraction unit 12 into voice to generate voice data. The method for converting text into voice is not particularly limited, and known text reading aloud (generation of synthetic voice from text) methods can be used.

[0017] The voice data and the captured image are associated and stored in the storage unit 10 as an image with voice.

[0018] The playback unit 14, in response to a request from the user terminal 2, causes the user terminal 2 to display the image with voice, plays back the voice data, and outputs the voice. The user can enjoy the voice together with the image.

[0019] A plurality of voice quality (tone color) data may be prepared so that a voice of a voice quality corresponding to the face of a person in the captured image (analog photo P) is output. In this case, as shown in FIG. 4, the server device 1 further has the functions of a face recognition unit 15 and a voice quality selection unit 16.

[0020] The face recognition unit 15 recognizes the face of a person in the captured image. The voice quality selection unit 16 selects an appropriate voice quality from the recognized face. For example, the voice quality selection unit 16 estimates the age, gender, build, etc. of the person from the face and selects the voice quality.

[0021] The voice data generation unit 13 converts the text into voice based on the selected voice quality to generate voice data. Thereby, a voice of a voice quality suitable for the person shown in the analog photo can be played back.

[0022] When the captured image includes the faces of a plurality of persons, voice data of voice qualities corresponding to the respective faces may be generated, synthesized, and simultaneously vocalized. Also, when the user selects a person from among the persons in the captured image 21 displayed on the user terminal 2, the voice data corresponding to the selected person may be played back.

[0023] It may be possible to save the voice quality settings for the person who generated the voice data so that the same voice quality can be used for the same person when using the service after the next time.

[0024] The voice quality selection unit 16 may randomly select the voice quality. In this case, a voice that does not match the person in the captured image (analog photo P) may be played, and the difference between the appearance of the person and the voice quality can be enjoyed.

[0025] A moving image in which the facial expression (mouth area) moves in accordance with the voice may be generated. In this case, as shown in FIG. 5, the server device 1 further has the function of the moving image generation unit 17.

[0026] The moving image generation unit 17 uses a known moving image generation method of lip-syncing a still image to the voice to generate a moving image in which the expression moves in accordance with the voice generated by the voice data generation unit 13 from the captured image. The moving image with voice is stored in the storage unit 10.

[0027] For example, when the user terminal 2 captures an analog photo P as shown in FIG. 2a, as shown in FIG. 6, a voice of "always a friend" is output, and a moving image in which the expression moves in accordance with the voice is displayed on the display of the user terminal 2.

[0028] The emotion of the person may be predicted based on the extracted text, and a moving image may be generated.

[0029] In the above embodiment, an example in which the analog photo P includes text has been described. However, for an analog photo P that does not include text, the user may select a favorite text from a plurality of texts prepared in advance, convert the selected text into voice, and generate an image with voice.

[0030] Although the present invention has been described in detail using specific embodiments, it is obvious to those skilled in the art that various changes can be made without departing from the intention and scope of the present invention.

Explanation of reference numerals

[0031] 1 Server device 2 User terminal

Claims

1. An image acquisition unit that acquires a captured image obtained by taking an analog photo from a user terminal; A text extraction unit that extracts text from the captured image; A face recognition unit that recognizes the faces of people in the captured image; A voice quality selection unit that selects a voice quality based on the recognized face; A voice data generation unit that converts the text into voice of the selected voice quality to generate voice data; A storage unit that stores the captured image and the voice data; A playback unit that displays the captured image on the user terminal and plays back the voice data; Comprising: A data generation system that, when the captured image includes the faces of a plurality of people, generates and plays back the voice data so that voices of voice qualities corresponding to the respective faces are simultaneously uttered.

2. The data generation system according to claim 1, further comprising a video generation unit that generates a moving image in which the expression of a person in the captured image moves in accordance with the voice data from the captured image.

Citation Information

Patent Citations

  • Electronic publish system, electronic diary, electronic album and recording medium

    JP1993003560A

  • Electronic album device and computer readable recording medium recording electronic album program

    JP2002190009A

  • Electronic album creation apparatus

    JP2006185223A

  • Electronic album creation device and program, and electronic album reproducer

    JP2006270144A

  • Method and apparatus for outputting voice

    JP2020008853A