Method, server and user terminal for providing audio based on captured image data
The method and system provide audio content for any picture book by analyzing photographed data and matching it with book index data, addressing the limitation of existing methods and enhancing user convenience by allowing accurate page-based audio playback.
Patent Information
- Application Number
- PCT/KR2024/011830
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-23
- Filing Date
- 2024-08-08
- Publication Date
- 2025-05-30
AI Technical Summary
Existing methods for providing audio to children, such as sound books and barcode recognition, are limited to specific fairy tale books and do not support existing general storybooks, making it difficult for caregivers in nuclear families to provide auditory content to children.
A method and system that utilize a server and user terminal to analyze photographed book data, extract text and image features, and match them with book index data to provide corresponding audio content, allowing for playback at the correct page position without requiring a built-in audio playback device.
Enables the provision of audio content for any picture book, enhancing user convenience by allowing audio playback to match the page content without manual flipping through the book.
Smart Images

Figure KR2024011830_30052025_PF_FP_ABST
Abstract
Description
Method for providing audio based on shooting data, server, and user terminal
[0001] The present invention relates to a method for providing audio based on captured data, a server, and a user terminal. Specifically, the present invention relates to a method for providing audio based on captured data, a server, and a user terminal that recognizes an image captured by a user and provides corresponding audio.
[0002]
[0003] This study was conducted as part of the 2024 Cultural Technology Research and Development Project of the Ministry of Culture, Sports and Tourism and the Korea Creative Content Agency (Project Name: Development of an Audiobook Playback System for Children Using AI-Based Storybook Recognition Technology, Project Number: RS-2024-00398248, Contribution Rate: 75%).
[0004] Reading has long been a recommended activity for all ages and genders. In particular, it was common for caregivers to read children's stories to foster cognitive, language, critical thinking, and social-emotional development. This allowed children to acquire visual information through storybooks while also acquiring auditory information through their caregivers' speech.
[0005] Today, with the rise of nuclear families and dual-income households, it's becoming increasingly difficult for caregivers to personally read stories to their children. Consequently, alternatives are emerging, such as storybooks with built-in audio playback devices, such as audio books, or methods that scan pre-printed barcodes to provide audio.
[0006] These alternatives have the limitation of not providing audio for existing picture books. Therefore, providing a method to provide audio for existing picture books is expected to greatly enhance user convenience.
[0007]
[0008] [Prior Art Literature]
[0009] [Patent Document]
[0010] Patent Registration No. 10-1393692
[0011]
[0012] The object of the present invention is to provide a method for providing audio based on photographed data, which derives at least one of text features and image features from an image of a book and transmits corresponding audio content to a user terminal.
[0013] Another object of the present invention is to provide an audio providing server that derives at least one of text features and image features from an image of a book and transmits audio content thereof to a user terminal.
[0014] The objectives of the present invention are not limited to those mentioned above. Other objectives and advantages of the present invention not mentioned above can be understood through the following description and will be more clearly understood through the embodiments of the present invention. Furthermore, it will be readily apparent that the objectives and advantages of the present invention can be realized by the means and combinations thereof set forth in the claims.
[0015]
[0016] According to some embodiments of the present invention for solving the above problem, a method for providing audio based on photographed data includes a step of analyzing original book data stored in a database of an audio providing server to generate book index data for text and pages, a step of receiving photographed image data of a portion of a book from a user terminal and generating analysis data for at least one of a text feature and an image feature of the photographed image data, a step of determining book index data matching the analysis data among the book index data as matching book data, and a step of determining a playback position of sound source data corresponding to the matching book data to generate playback data, and a step of providing the playback data to the user terminal.
[0017] Additionally, the book index data may include text feature data for text included in the book original data, image feature data for images included in the book original data, and page information related to the text feature data and the image feature data.
[0018] Additionally, the above-mentioned book index data may further include timestamp data for the playback position of the above-mentioned sound source data.
[0019] In addition, the timestamp data can be generated through a step of generating subtitle data for a script of the sound source data and a step of comparing the text feature data and the subtitle data to generate the timestamp data for a playback position of the sound source data corresponding to the text feature data.
[0020] Additionally, the step of generating the analysis data may include at least one of a step of generating text feature data for text included in the photographed image data and a step of generating image feature data for image features included in the photographed image data.
[0021] In addition, the step of determining the matching book data may include at least one of a step of calculating a first similarity between the text feature data and the book index data, a step of calculating a second similarity between the image feature data and the book index data, a step of calculating a comprehensive similarity between each of the book index data based on at least one of the first and second similarities, and a step of determining the book index data having the highest comprehensive similarity as the matching book data.
[0022] In addition, the step of determining the matching book data may include a step of determining the page with the highest overall similarity as the matching page and a step of determining book index data including the matching page as the matching book data.
[0023] In addition, the step of providing to the user terminal may include a step of providing sound source data corresponding to the matching book data from the database, a step of determining a timestamp corresponding to a matching page matching the analysis data from the sound source data provided from the database as a playback position, a step of generating the playback data including the sound source data and the playback position, and a step of providing the playback data to the user terminal.
[0024] An audio providing server according to some embodiments of the present invention for solving the above problem includes a database for storing book index data and audio data for a book, a communication module for receiving photographed image data of a portion of a book from a user terminal, an analysis module for generating analysis data on text and image features of the photographed image data, a search module for generating matching book data by matching the book index data and the analysis data, and an audio processing module for generating playback data by determining a playback position of audio data corresponding to the matching book data.
[0025] An audio providing server according to some embodiments of the present invention for solving the above problem includes a database for storing book index data and sound source data for a book, a shooting module for shooting a portion of a book to generate shot image data, an analysis module for generating analysis data for at least one of text features and image features of the shot image data, a search module for matching the book index data and the analysis data to generate matching book data, and an audio processing module for determining a playback position of sound source data corresponding to the matching book data to generate playback data, and providing the playback data to the user terminal.
[0026]
[0027] The method, server, and user terminal of the present invention for providing audio based on photographed images can extract at least one of text features and image features from a photographed image of a book and match them with book data stored in a database. This allows audio content for picture books without built-in audio playback devices to be provided to users.
[0028] Additionally, the audio provision method, server, and user terminal of the present invention based on captured data can designate the playback position of audio content to correspond to the captured page of the book. This eliminates the need for the user to flip through the book to match the audio content, thereby enhancing user convenience.
[0029] In addition to the above-described contents, the specific effects of the present invention are described together with the specific matters for carrying out the invention below.
[0030]
[0031] FIG. 1 is a conceptual diagram illustrating a system for providing audio based on shooting data according to some embodiments of the present invention.
[0032] FIG. 2 is a block diagram for explaining in detail the audio providing server of FIG. 1.
[0033] FIG. 3 is a flowchart illustrating a method for providing audio based on shooting data according to some embodiments of the present invention.
[0034] Figure 4 is a block diagram for explaining in detail the analysis module of Figure 2.
[0035] Figure 5 is a flowchart for explaining step S300 of Figure 3 in detail.
[0036] Figure 6 is a conceptual diagram schematically explaining a neural network applied to the analysis module of Figure 4.
[0037] Figure 7 is a block diagram for explaining in detail the search module of Figure 2.
[0038] Figure 8 is a flowchart for explaining step S400 of Figure 3 in detail.
[0039] Figure 9 is a flowchart for explaining step S500 of Figure 3 in detail.
[0040] Figure 10 is an example diagram for explaining the image area analyzed by the analysis module of Figure 2.
[0041] FIG. 11 is a block diagram for explaining in detail an audio providing server performing step S100 of FIG. 3.
[0042] Figure 12 is a block diagram for explaining in detail the analysis module of Figure 11.
[0043] Figure 13 is a flowchart for explaining step S100 of Figure 3 in detail.
[0044] Figure 14 is a diagram of an example of book index data (BID) corresponding to Figure 10.
[0045] Figure 15 is a block diagram for explaining in detail the text analysis module of Figure 12.
[0046] Figure 16 is a flowchart for explaining steps S130 and S140 of Figure 13 in detail.
[0047] FIG. 17 is a block diagram illustrating a user terminal that performs an audio providing method according to some embodiments of the present invention.
[0048] FIG. 18 is a diagram for explaining the hardware configuration of an audio providing server or system that performs a method for providing audio based on shooting data according to some embodiments of the present invention.
[0049]
[0050] The terms and words used in this specification and claims should not be interpreted based on their general or dictionary meanings. In accordance with the principle that inventors can define the concepts of terms and words to best describe their inventions, they should be interpreted in a way that is consistent with the technical concept of the present invention. Furthermore, the embodiments described in this specification and the configurations depicted in the drawings are merely examples of how the present invention can be realized and do not fully represent the technical concept of the present invention. Therefore, it should be understood that various equivalents, modifications, and applicable examples may exist as of the time of filing.
[0051] The terms first, second, A, B, etc. used in this specification and claims may be used to describe various components, but the components should not be limited by these terms. These terms are used only for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be referred to as the second component, and similarly, the second component may also be referred to as the first component. The term "and / or" includes any combination of a plurality of related listed items or any item among a plurality of related listed items.
[0052] The terminology used in this specification and claims is for the purpose of describing specific embodiments only and is not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly dictates otherwise. It should be understood that terms such as "comprise" or "have" in this application do not preclude the presence or addition of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification.
[0053] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.
[0054] Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless expressly defined in this application.
[0055] In addition, each configuration, process, procedure or method included in each embodiment of the present invention may be shared within a scope that is not technically inconsistent with each other.
[0056]
[0057] Hereinafter, with reference to FIGS. 1 to 18, a method for providing audio based on shooting data, a server, and a user terminal according to some embodiments of the present invention will be described.
[0058] FIG. 1 is a conceptual diagram illustrating a system for providing audio based on shooting data according to some embodiments of the present invention.
[0059] Referring to FIG. 1, a system for providing audio based on shooting data according to some embodiments of the present invention includes an audio providing server (100), a user terminal (200), and a communication network (300). Here, the audio providing server (100) operates in connection with the user terminal (200) via the communication network (300) and can transmit and receive data.
[0060] The audio provision server (100) can be linked with an audio provision program (e.g., a mobile application) running on a user terminal (200) and can be an entity that provides various services including provision of audio content.
[0061] In the present invention, the audio providing server (100) may operate as an executor of a method for providing audio based on shooting data. However, the present invention is not limited thereto, and each step included in the method for providing audio based on shooting data of the present invention may be performed separately in some steps in the audio providing server (100) and the user terminal (200), or all steps may be performed in the user terminal (200) rather than the audio providing server (100). In other words, the audio providing method of the present invention is not limited to a specific executor that performs it. However, for the convenience of explanation, the audio providing server (100) as an executor of the method for providing audio based on shooting data of the present invention will be described as an example with reference to FIG. 1 to FIG. 16, and the user terminal (201) as an executor of the method for providing audio based on shooting data of the present invention will be described as an example with reference to FIG. 17.
[0062] The user terminal (200) refers to a terminal of a user who uses the service provided by the audio provision server (100). The user can use the user terminal (200) to take a picture of a portion of a book and use audio content corresponding to the page of the book taken. As a specific example, the user can transmit an image containing text or pictures of a book to the audio provision server (100) through the user terminal (200) and retrieve book page information for the image, thereby receiving audio content related to the image containing text or pictures of the book provided to the audio provision server (100).
[0063] Meanwhile, the communication network (300) plays a role of connecting the audio providing server (100) and the user terminal (200). That is, the communication network (300) refers to a communication network that provides a connection path so that the user terminal (200) can transmit and receive data after connecting to the audio providing server (100). The communication network (300) may include wired networks such as LANs (Local Area Networks), WANs (Wide Area Networks), MANs (Metropolitan Area Networks), and ISDNs (Integrated Service Digital Networks), or wireless networks such as wireless LANs, CDMA, Bluetooth, and satellite communication, but the scope of the present invention is not limited thereto.
[0064]
[0065] FIG. 2 is a block diagram for explaining in detail the audio provision server (100) of FIG. 1.
[0066] Referring to FIG. 2, the audio provision server (100) may include a communication module (110), an analysis module (130), a search module (150), an audio processing module (170), and a database (DB).
[0067] The communication module (110) can communicate with an external device or a user terminal (200) via a communication network (300). The communication module (110) can receive captured image data (OID) from the user terminal (200). In addition, the communication module (110) can transmit reproduction data (AFP) corresponding to the captured image data (OID) to the user terminal (200).
[0068] The analysis module (130) can analyze captured image data (OID) and generate analysis data (AD) therefor. Specifically, the analysis module (130) can perform analysis on at least one of the text features and image features of the captured image data (OID). A detailed description of the analysis module will be described later with reference to FIGS. 4 and 5 .
[0069] The search module (150) can compare the book index data (BID) received from the database (DB) with the analysis data (AD). Through this, the search module (150) can determine the book index data (BID) that matches the analysis data (AD) as the matching book data (MBD). A detailed description of the search module will be described later with reference to FIGS. 6 and 7.
[0070] The database (DB) can store original book data, including text and image data, and corresponding audio data (OAD). Additionally, the database (DB) can store book index data (BID) for the book's text and pages.
[0071] The audio processing module (170) can receive sound source data (OAD) corresponding to matching book data (MBD) from a database (DB). Subsequently, the audio processing module (170) can determine the playback position of the sound source data (OAD) based on the matching book data (MBD) and generate playback data (AFP). The audio processing module (170) can provide the generated playback data (AFP) to the user terminal (200) via the communication module (110).
[0072]
[0073] Figure 3 is a flowchart illustrating a method for providing audio based on captured data according to some embodiments of the present invention. Parts that overlap with the above description are omitted or simplified.
[0074] Referring to FIGS. 2 and 3, the audio provision server (100) analyzes the original book data to generate book index data for text and pages (S100). Step S100 will be described later with reference to FIGS. 11 to 16.
[0075] Next, the audio provision server (100) receives captured image data from the user terminal (200) (S200).
[0076] Specifically, the audio provision server (100) can receive captured image data (OID) from the user terminal (200) via the communication module (110). At this time, the captured image data (OID) is an image of a portion of a book and may include at least one of the text and image of the book. In addition, the communication module (110) can transmit the captured image data (OID) to the analysis module (130).
[0077] Next, the audio provision server (100) generates analysis data for at least one of the text features and image features of the captured image data (S300).
[0078] Specifically, the analysis module (130) can recognize at least one of text features and image features included in the captured image data (OID) to generate analysis data (AD). In other words, the analysis data (AD) can include data on at least one of text features (e.g., features for the main text of a book, features for pages, etc.) and image features (features for rabbit images, features for tree images, etc.) included in the captured image data (OID).
[0079] At this time, the analysis module (130) may utilize various known deep learning models. Typically, a deep learning model may comprise a neural network composed of multiple layers, each layer connected to multiple neural networks. A detailed description of the analysis module (130) will be provided below with reference to FIGS. 4 to 6.
[0080] Next, the audio provision server (100) determines the book index data that matches the analysis data among the book index data as the matching book data (S400).
[0081] Specifically, the search module (150) may be provided with book index data (BID) stored in a database (DB). Here, the book index data (BID) may include data on at least one of text features, image features, pages, and audio playback positions of multiple books.
[0082] Additionally, the search module (150) can receive analysis data (AD) from the analysis module (130). The search module (150) can determine the book index data (BID) matching the analysis data (AD) as matching book data (MBD). Here, the matching book data (MBD) can additionally include matching page data for the matching page corresponding to the photographed image data (OID). A detailed description of the search module (150) will be described later with reference to FIGS. 7 and 8.
[0083] Next, the audio provision server (100) determines the playback position of the sound source data corresponding to the matching book data, generates playback data, and provides it to the user terminal (200) (S500).
[0084] Specifically, the audio processing module (170) can receive sound source data (OAD) corresponding to matching book data (MBD) from a database (DB). The audio processing module (170) can determine a timestamp corresponding to matching page data in the sound source data (OAD) as a playback position. Subsequently, the audio processing module (170) can generate playback data (AFP) including the sound source data (OAD) and the playback position, and provide the same to the user terminal (200) via the communication module (110).
[0085]
[0086] Fig. 4 is a block diagram for explaining in detail the analysis module of Fig. 2, and Fig. 5 is a flowchart for explaining in detail step S300 of Fig. 3. Fig. 6 is a conceptual diagram for schematically explaining a neural network applied to the analysis module of Fig. 4.
[0087] Referring to FIG. 4, the analysis module (130) may include a text analysis module (131) and an image analysis module (133).
[0088] The text analysis module (131) can analyze text features included in captured image data (OID) to generate text feature data (TFD). Meanwhile, the image analysis module (133) can analyze image features included in captured image data (OID) to generate image feature data (IFD).
[0089] Referring further to FIG. 5, the analysis module (130) generates text feature data for characters included in the captured image data (S310).
[0090] Specifically, the text analysis module (131) can generate text feature data (TFD) for text included in the captured image data (OID). At this time, the text feature data (TFD) can be included in the analysis data (AD). For example, the text feature data (TFD) can include at least one of a text feature vector, a tokenized word, and a sentence.
[0091] The analysis module (130) generates image feature data for image features included in the captured image data (S320).
[0092] Specifically, the image analysis module (133) can generate image feature data (IFD) for image features included in the captured image data (OID). At this time, the image feature data (IFD) can be included in the analysis data (AD). In addition, the image feature data (IFD) can include at least one of an image feature vector and a label for the image feature. However, the operation of the analysis module (130) is not limited to the order of steps S310 and S320 illustrated in FIG. 5, and the embodiments are not limited thereto. For example, the analysis module (130) may operate by omitting either step S310 or S320, may perform other operations in addition to steps S310 and S320, may operate in the opposite order to the illustrated order, and steps S310 and S320 may operate in parallel with each other. That is, FIG. 5 is only intended to explain the operations performed by the analysis module (130) and does not represent their time-series relationship.
[0093] Referring to FIG. 6, a neural network (NN) constituting a text analysis module (131) or an image analysis module (133) may include an input layer (Input) including input nodes, an output layer (Output) including output nodes, and M hidden layers arranged between the input layer (Input) and the output layer (Output).
[0094] Specifically, in the case of the first neural network constituting the text analysis module (131), captured image data (OID) may be input to the input node, and text feature data (TFD) may be output from the output node.
[0095] Additionally, in the case of the second neural network constituting the image analysis module (133), captured image data (OID) may be applied to the input node, and image feature data (IFD) may be output from the output node.
[0096] For example, a deep learning model such as a CNN (Convolutional Neural Network), an RNN (Recurrent Neural Network), a DBN (Deep Belief Network), a GNN (Graph Neural Network), or a Transformer model may be used in the neural network that constitutes the text analysis module (131) or the image analysis module (133).
[0097] For example, the text analysis module (131) may utilize a pre-trained deep learning model to perform optical character recognition (OCR). Meanwhile, the image analysis module (133) may utilize the VGG16 model, a type of CNN. However, the present embodiment is not limited thereto.
[0098] Here, weights can be assigned to the edges connecting the nodes of each layer that constitutes the deep learning model. These weights or edges can be added, removed, or updated during the learning process. Therefore, the weights of the nodes and edges between the k input nodes and i output nodes can be updated during the learning process.
[0099] The text analysis module (131) or the image analysis module (133) may be pre-trained before analyzing the captured image data (OID). Before the text analysis module (131) or the image analysis module (133) performs training, all nodes and edges may be set to initial values. However, when information is input cumulatively, the weights of the nodes and edges are changed, and in this process, a matching may be made between the parameters input as training factors (e.g., original book data) and the values assigned to the output nodes (e.g., text feature data or image feature data).
[0100]
[0101] FIG. 7 is a block diagram for explaining in detail the search module of FIG. 2, and FIG. 8 is a flowchart for explaining in detail step S400 of FIG. 3.
[0102] Referring to FIG. 7, the search module (150) may include a text similarity calculation unit (151), an image similarity calculation unit (153), a comprehensive similarity calculation unit (155), and a matching book selection unit (157).
[0103] The text similarity calculation unit (151) can generate a first similarity (Sml1) based on text feature data (TFD) and book index data (BID). Meanwhile, the image similarity calculation unit (153) can generate a second similarity (Sml2) based on image feature data (IFD) and book index data (BID). The comprehensive similarity calculation unit (155) can generate a comprehensive similarity (Sml_T) and matching page data (MPD) based on the first similarity (Sml1) and the second similarity (Sml2). Subsequently, the matching book selection unit (157) can generate matching book data (MBD) based on the comprehensive similarity (Sml_T) and matching page data (MPD).
[0104] Referring further to FIG. 8, the search module (150) calculates a first similarity between the text feature data and the book index data (S410).
[0105] The text similarity calculation unit (151) can receive and compare text feature data (TFD) and book index data (BID). For example, the text similarity calculation unit (151) can compare the text of each of a plurality of books included in the book index data (BID) with the text feature data (TFD) to calculate a first similarity (Sml1) for each book. Subsequently, the text similarity calculation unit (151) can transmit the calculated first similarity (Sml1) to the comprehensive similarity calculation unit (155).
[0106] At this time, the first similarity (Sml1) can be calculated for each page of each book index data (BID). For example, the first similarity (Sml1) can be calculated using an inverted index structure and a scoring algorithm.
[0107] Next, the search module (150) calculates a second similarity between the image feature data and the book index data (S420).
[0108] Specifically, the image similarity calculation unit (153) can receive image feature data (IFD) and book index data (BID) and compare them with each other. For example, the image similarity calculation unit (153) can compare the image features of each of a plurality of books included in the book index data (BID) with the image feature data (IFD) to calculate a second similarity (Sml2) for each book. Subsequently, the image similarity calculation unit (153) can transmit the calculated second similarity (Sml2) to the comprehensive similarity calculation unit (155).
[0109] At this time, the second similarity (Sml2) can be calculated for each page of each book index data (BID). For example, the second similarity (Sml2) can be calculated using an ANN algorithm (Approximate Nearest Neighbor algorithm).
[0110] Next, the search module (150) calculates a comprehensive similarity for each book index data based on at least one of the first similarity and the second similarity (S430).
[0111] Specifically, the comprehensive similarity calculation unit (155) can calculate the comprehensive similarity (Sml_T) based on at least one of the first similarity (Sml1) and the second similarity (Sml2).
[0112] For example, if the captured image data (OID) contains only text, text feature data (TFD) may be generated, but image feature data (IFD) may not be generated. Accordingly, the first similarity (Sml1) may be calculated from the text feature data (TFD), but the second similarity (Sml2) may not be calculated. In this case, the comprehensive similarity (Sml_T) may be calculated using only the first similarity (Sml1).
[0113] As another example, if the captured image data (OID) only contains images, image feature data (IFD) may be generated, but text feature data (TFD) may not be generated. Accordingly, the second similarity (Sml2) may be calculated from the image feature data (IFD), but the first similarity (Sml1) may not be calculated. In this case, the comprehensive similarity (Sml_T) may be calculated using only the second similarity (Sml2).
[0114] As another example, if the captured image data (OID) contains both text and images, both text feature data (TFD) and image feature data (IFD) can be generated. Accordingly, the first similarity (Sml1) can be derived from the text feature data (TFD), and the second similarity (Sml2) can be derived from the image feature data (IFD). In this case, the comprehensive similarity (Sml_T) can be derived based on the first and second similarities (Sml1, Sml2).
[0115] In other words, the comprehensive similarity (Sml_T) may be a value that comprehensively evaluates the similarity between the photographed image data (OID) and the book index data (BID) by comparing at least one of the text features and the image features.
[0116] At this time, the similarity can be calculated using various known methods. For example, the text similarity calculation unit (151), the image similarity calculation unit (153), and the comprehensive similarity calculation unit (153) can calculate the similarity using mean squared difference similarity, cosine similarity, Minkowski distance, Jaccard similarity, Pearson similarity, etc.
[0117] Next, the search module (150) determines the page with the highest overall similarity as the matching page (S440).
[0118] Specifically, the comprehensive similarity calculation unit (155) can calculate the comprehensive similarity (Sml_T) for each page of the book index data (BID). Here, the comprehensive similarity calculation unit (155) can determine the page with the highest comprehensive similarity (Sml_T) as the matching page. Subsequently, the comprehensive similarity calculation unit (155) can generate matching page data (MPD) including matching pages for each book of the book index data (BID) and transmit this to the matching book selection unit (157).
[0119] Next, the search module (150) determines the book index data including the matching page as the matching book data (S450).
[0120] Specifically, the matching book selection unit (157) can select matching books based on the comprehensive similarity (Sml_T) and matching page data (MPD). Subsequently, the matching book selection unit (157) can generate book index data (BID) including matching pages and matching book data (MPD) including matching page data (MPD).
[0121]
[0122] Figure 9 is a flowchart for explaining step S500 of Figure 3 in detail. Any content that overlaps with the above is omitted or simplified.
[0123] Referring to FIG. 2 and FIG. 9, the audio processing module (170) receives sound source data corresponding to matching book data from the database (S510).
[0124] Specifically, the audio processing module (170) can receive matching book data (MBD) from the search module (150). In addition, the audio processing module (170) can receive sound source data (OAD) corresponding to the matching book data (MBD) from a database (DB).
[0125] Next, the audio processing module (170) determines the timestamp corresponding to the matching page from the sound source data provided from the database as the playback position (S520).
[0126] The audio processing module (170) can check the matching page included in the matching book data (MBD) and the timestamp corresponding to the matching page. The audio processing module (170) can determine the timestamp corresponding to the matching page as the playback position.
[0127] Next, the audio processing module (170) generates playback data (AFP) including sound source data (OAD) and playback position (S530). Next, the audio processing module (170) provides the playback data (AFP) to the user terminal (200) (S540).
[0128] In some embodiments of the present invention, the book index data (BID) and matching book data (MBD) may not include timestamp data (TD). In this case, the audio processing module (170) may compare the script of the audio data (OAD) with the text feature data (TFD) of the matching book data (MBD) to derive a timestamp corresponding to the matching page. The audio processing module (170) may determine the derived timestamp as a playback position and generate playback data (AFP) including the audio data (OAD) and the playback position.
[0129]
[0130] Figure 10 is an example diagram illustrating the image area analyzed by the analysis module of Figure 2. Parts overlapping with the above-described content are omitted or simplified.
[0131] Referring to FIG. 10, a user can take a picture of a portion of a book and generate picture image data (OID) for the picture image area (OIA).
[0132] The analysis module (130) can distinguish between a text extraction area (TEA), which is an area containing characters (text) within the captured image area (OIA), and an image feature extraction area (IEA), which is an area containing image features.
[0133] The text analysis module (131) can identify characters in a text extraction area (TEA) to generate text feature data (TFD). For example, the text feature data (TFD) may include at least one of a text feature vector, tokenized word, and sentence for the text "The tortoise and the hare started a race. The tortoise crawled slowly on its short legs." contained in the text extraction area (TEA).
[0134] Meanwhile, the image analysis module (133) can generate image feature data (IFD) by extracting image features from the image feature extraction area (IEA). For example, the image feature data (IFD) can include an image feature vector for at least one of a turtle and a tree, which are images included in the image feature extraction area (IEA).
[0135] As another example, the captured image data (OID) may only contain regions containing characters, and may not contain regions containing image features. In this case, the text analysis module (131) can identify characters in the text extraction area (TEA) to generate text feature data (TFD). On the other hand, since there are no regions containing image features, image feature data (IFD) may not be generated.
[0136] Conversely, the captured image data (OID) may only contain regions containing image features, and may not contain regions containing text. In this case, the image analysis module (133) may extract image features from the image extraction area (IEA) to generate image feature data (IFD). On the other hand, since there is no region containing text, text feature data (IFD) may not be generated.
[0137]
[0138] Fig. 11 is a block diagram for explaining in detail the audio provision server performing step S100 of Fig. 3, and Fig. 12 is a block diagram for explaining in detail the analysis module of Fig. 11. Fig. 13 is a flowchart for explaining in detail step S100 of Fig. 3. Parts overlapping with the above contents are omitted or simplified.
[0139] Referring to Figure 11, the analysis module (130) can receive original book data (OBD) and audio data (OAD) from a database (DB). Here, the original book data (OBD) may include text and image data for the entire book. For example, the original book data (OBD) may be an e-book file in PDF format.
[0140] The analysis module (130) can analyze the original book data (OBD) and the audio data (OAD) to generate book index data (BID). Here, the book index data (BID) can include data regarding at least one of the text, image, page, and timestamp of the audio data of the original book data.
[0141] Specifically, referring to FIG. 12, the analysis module (130) may include a text analysis module (131), an image analysis module (133), and a mapping module (135).
[0142] Referring further to FIG. 13, the analysis module (130) generates text feature data for text included in the original book data (S110).
[0143] In detail, the text analysis module (131) can recognize text included in the original book data (OBD) and generate text feature data (TFD). A detailed description of the text analysis module (131) will be described later with reference to FIGS. 15 and 16.
[0144] Next, the analysis module (130) generates image feature data for the images included in the original book data (S120).
[0145] In detail, the image analysis module (133) can generate image feature data (IFD) for images included in the original book data (OBD). At this time, the image analysis module (133) can operate as described in FIGS. 4 to 6.
[0146] Next, the analysis module (130) generates page information related to text feature data and image feature data (S130).
[0147] Specifically, the mapping module (135) can receive text feature data (TFD) and image feature data (IFD). The mapping module (135) can generate page information (PD of FIG. 16) related to the text feature data (TFD) and the image feature data (IFD). Here, the page information (PD) can include information about a corresponding page within a book in the text feature data (TFD) and information about a corresponding page within a book in the image feature data (IFD). For example, the page information can include information that the text feature data (TFD) of 'The rabbit hopped with long legs' and the image feature data (IFD) of 'rabbit, tree, road' are arranged on 'page 7'.
[0148] Next, the analysis module (130) generates timestamp data for the playback position of the sound source data corresponding to the original book data (S140).
[0149] Specifically, the text analysis module (131) can generate timestamp data (TD) for the playback position of audio data (OAD). Here, the timestamp data (TD) may be data regarding the mapping relationship between the script of the audio data (OAD) and the playback position.
[0150] Next, the analysis module (130) generates book index data by mapping text feature data, image feature data, page information, and timestamp data (S150).
[0151] Specifically, the mapping module (135) can generate book index data (BID) by mapping text feature data (TFD), image feature data (IFD), page information (PD), and timestamp data (TD). In other words, the book index data (BID) can include text feature data (TFD), image feature data (IFD), and timestamp data (TD) corresponding to each page of the book. For an exemplary description of the book index data (BID), refer to FIG. 14.
[0152]
[0153] Figure 14 illustrates an example of book index data (BID) corresponding to Figure 10.
[0154] Referring to FIG. 14, the text feature data (TFD) of the first book index data (BID_1) may include "The hare and the tortoise started a race" and "The tortoise crawled slowly on its short legs." Furthermore, the image feature data (IFD) of the first book index data (BID_1) may include "turtle," "tree," and "road." The page information (PD) of the first book index data (BID_1) may be "Page 6," and the timestamp data (TD) may be "00:07:23."
[0155] Similarly, the text feature data (TFD) of the second book index data (BID_2) may include “The rabbit hopped along with its long legs.” Furthermore, the image feature data (IFD) of the second book index data (BID_2) may include “feature vectors for rabbit images,” “feature vectors for tree images,” and “feature vectors for road images.” The page information (PD) of the second book index data (BID_2) may be “page 7,” and the timestamp data (TD) may be “00:07:51.”
[0156] In Figure 14, the text feature data (TFD) is described using a sentence in Korean as an example for convenience of explanation. However, the present embodiment is not limited thereto, and the text feature data (TFD) may be at least one of a text feature vector and a tokenized word.
[0157]
[0158] Fig. 15 is a block diagram for explaining in detail the text analysis module of Fig. 12, and Fig. 16 is a flowchart for explaining in detail steps S130 and S140 of Fig. 13.
[0159] Referring to FIGS. 15 and 16, the text analysis module (131) may include a voice recognition unit (131_1), a character recognition unit (131_2), and a timestamp generation unit (131_3).
[0160] First, the text analysis module (131) generates subtitle data for the script of the sound source data (S141).
[0161] Specifically, the voice recognition unit (131_1) can receive audio data (OAD) and recognize the voice of the audio data (OAD). Here, the audio data (OAD) is data of a fairy tale being read aloud, and may include an audio book. Accordingly, the script of the audio data (OAD) may be identical to the main text of the fairy tale.
[0162] The voice recognition unit (131_1) can generate subtitle data (SD) in text form based on the script of the voice recognized from the audio source data (OAD). Here, the voice recognition unit (131_1) can be configured as a neural network. The voice recognition unit (131_1) can be configured and operated in a manner similar to the neural network described above in FIG. 6.
[0163] In detail, in the case of the third neural network constituting the voice recognition unit (131_1), audio data (OAD) in the form of voice can be input to the input node, and subtitle data (SD) in the form of text can be output from the output node.
[0164] Next, the text analysis module (131) compares the text feature data and the subtitle data to generate timestamp data for the playback position of the sound source data corresponding to the text feature data (S143).
[0165] Specifically, the character recognition unit (131_2) can recognize text included in the original book data (OBD) and generate text feature data (TFD). Subsequently, the character recognition unit (131_2) can transmit the generated text feature data (TFD) to the timestamp generation unit (131_3).
[0166] The timestamp generation unit (131_3) can compare subtitle data (SD) and text feature data (TFD) to derive a correspondence between the text feature data (TFD) and the playback position of the audio source data (OAD). The timestamp generation unit (131_3) can generate timestamp data (TD) including the correspondence between the text feature data (TFD) and the playback position.
[0167]
[0168] Figure 17 is a block diagram illustrating a user terminal performing an audio provision method according to some embodiments of the present invention. Parts that overlap with the above description are omitted or simplified.
[0169] The user terminal (201) of Fig. 17 may include a shooting module (190), an analysis module (130), a search module (150), an audio processing module (170), and a database (DB). The user terminal (201) refers to a user's terminal that directly performs the service provided by the audio provision server (100) described above.
[0170] The photographing module (190) can directly photograph a portion of a book to generate photographed image data (OID). Subsequently, the photographing module (190) can transmit the photographed image data (OID) to the analysis module (130). As described above, the analysis module (130) can generate analysis data (AD), and the search module (150) can generate matching book data (MBD). Subsequently, the audio processing module (170) can provide playback data (AFP) through an external output device (not shown) without going through the communication module (110) to reproduce audio.
[0171]
[0172] The method, server, and user terminal for providing audio based on photographed data according to the present embodiments can extract at least one of text features and image features from photographed image data of a portion of a book and match them with book index data stored in a database. This allows audio content for picture books without built-in audio playback devices to be provided to users.
[0173] Furthermore, the present invention can determine the playback position of audio content to correspond to the page of captured image data, thereby allowing audio content to be played from that playback position. This eliminates the need for users to flip through the book to find the audio content, thereby enhancing user convenience.
[0174]
[0175] FIG. 18 is a diagram for explaining the hardware configuration of an audio providing server, user terminal, or system that performs a method for providing audio based on shooting data according to some embodiments of the present invention.
[0176] Referring to FIG. 18, an audio providing server (100) or a user terminal (200, 201) according to some embodiments of the present invention may be implemented as an electronic device (1000). The electronic device (1000) may include a processor (1010), an input / output device (1020, I / O), a memory (1030, memory), an interface (1040), a storage (1050, storage), and a bus (1060, bus). The processor (1010), the input / output device (1020), the memory (1030), the interface (1040), and / or the storage (1050) may be coupled to each other via a bus (1060). The bus (1060) corresponds to a path through which data is transferred.
[0177] Specifically, the processor (1010) may include at least one of a Central Processing Unit (CPU), a Micro Processor Unit (MPU), a Micro Controller Unit (MCU), a Graphic Processing Unit (GPU), a microprocessor, a digital signal processor, a microcontroller, an application processor (AP), and logic elements capable of performing functions similar thereto.
[0178] The input / output device (1020) may include at least one of a keypad, a keyboard, a touchscreen, and a display device.
[0179] The memory (1030) can load data and / or programs, etc. At this time, the memory (1030) is an operating memory for improving the operation of the processor (1010) and may include high-speed DRAM and / or SRAM. The memory (1030) may include one or more volatile memory devices such as DDR SDRAM (Double Data Rate Static DRAM) and SDR SDRAM (Single Data Rate SDRAM) and / or one or more non-volatile memory devices such as EEPROM (Electrically Erasable Programmable ROM) and flash memory.
[0180] The interface (1040) may perform a function of transmitting data to or receiving data from a communication network. The interface (1040) may be wired or wireless. For example, the interface (1040) may include an antenna or a wired or wireless transceiver.
[0181] Storage (1050) can store and preserve data and / or programs. Storage (1050) may include one or more non-volatile memory devices, such as a solid state drive (SSD), a hard drive, or flash memory. In the present invention, storage (1050) may store a computer program comprising instructions for performing an image-based account transfer service method.
[0182] Alternatively, the audio provision server (100) and user terminals (200, 201) according to embodiments of the present invention may each be a system formed by connecting multiple electronic devices (1000) to each other via a network. In this case, each module or combination of modules may be implemented as an electronic device (1000). However, the present embodiment is not limited thereto.
[0183] Additionally, the audio providing server (100) may be implemented as at least one of a workstation, a data center, an internet data center (IDC), a direct attached storage (DAS) system, a storage area network (SAN) system, a network attached storage (NAS) system, and a redundant array of inexpensive disks (RAID) system, but the present embodiment is not limited thereto.
[0184] Additionally, the audio provision server (100) can transmit data via a network using a user terminal (200). The network may include a network based on wired Internet technology, wireless Internet technology, and short-range communication technology. For example, the wired Internet technology may include at least one of a local area network (LAN) and a wide area network (WAN).
[0185] The wireless Internet technology may include, for example, at least one of Wireless LAN (WLAN), Digital Living Network Alliance (DMNA), Wireless Broadband (Wibro), World Interoperability for Microwave Access (Wimax), High Speed Downlink Packet Access (HSDPA), High Speed Uplink Packet Access (HSUPA), IEEE 802.16, Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), Wireless Mobile Broadband Service (WMBS), and 5G NR (New Radio) technologies. However, the present embodiment is not limited thereto.
[0186] Short-range communication technologies may include, for example, at least one of Bluetooth, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra-Wideband (UWB), ZigBee, Near Field Communication (NFC), Ultra Sound Communication (USC), Visible Light Communication (VLC), Wi-Fi, Wi-Fi Direct, and 5G NR (New Radio). However, the present embodiment is not limited thereto.
[0187] The audio providing server (100) communicating through a network may comply with technical standards and standard communication methods for mobile communication. For example, the standard communication method may include at least one of GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), CDMA2000 (Code Division Multi Access 2000), EV-DO (Enhanced Voice-Data Optimized or Enhanced Voice-Data Only), WCDMA (Wideband CDMA), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTEA (Long Term Evolution-Advanced), and 5G NR (New Radio). However, the present embodiment is not limited thereto.
[0188]
[0189] The above description is merely an example of the technical idea of the present embodiment, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential characteristics of the present embodiment. Therefore, the present embodiments are not intended to limit the technical idea of the present embodiment, but rather to explain it, and the scope of the technical idea of the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment should be interpreted by the claims below, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of rights of the present embodiment.
Claims
1. In a method for providing audio based on shooting data performed on an audio providing server linked to a user terminal, A step of analyzing original book data stored in the database of the above audio providing server to generate book index data for text and pages; A step of receiving a photographed image data of a portion of a book from the user terminal, and generating analysis data for at least one of a text feature and an image feature of the photographed image data; A step of determining the book index data that matches the analysis data among the above book index data as the matching book data; and A step of determining a playback position of sound data corresponding to the above matching book data to generate playback data and providing the playback data to the user terminal is included. Method for providing audio based on captured data.
2. In paragraph 1, The above book index data is, Text feature data for the text included in the above book source data; Image feature data for images included in the above book original data; and Contains page information related to the above text feature data and the above image feature data. Method for providing audio based on captured data.
3. In paragraph 2, The above book index data further includes timestamp data for the playback position of the above audio data. Method for providing audio based on captured data.
4. In paragraph 3, The above timestamp data is, A step for generating subtitle data for a script of the above sound source data, and The step of generating the timestamp data for the playback position of the sound source data corresponding to the text feature data by comparing the text feature data and the subtitle data is performed. Method for providing audio based on captured data.
5. In paragraph 1, The steps for generating the above analysis data are: A step of generating text feature data for text included in the above photographed image data; and At least one step of generating image feature data for image features included in the above-mentioned photographed image data Method for providing audio based on captured data.
6. In paragraph 5, The step of determining the above matching book data is: At least one of a step of calculating a first similarity with the text feature data for the above book index data and a step of calculating a second similarity with the image feature data for the above book index data; A step of calculating a comprehensive similarity for each of the book index data based on at least one of the first and second similarities; and A step of determining the book index data with the highest comprehensive similarity as the matching book data is included. Method for providing audio based on captured data.
7. In paragraph 6, The step of deciding with the above matching book data is: A step of determining the page with the highest overall similarity as the matching page; and A step of determining the book index data including the above matching page as the above matching book data. Method for providing audio based on captured data.
8. In paragraph 1, The steps provided to the above user terminal are: A step of providing sound source data corresponding to the matching book data in the above database; A step of determining a timestamp corresponding to a matching page matching the analysis data from the sound source data provided from the above database as a playback position; A step of generating the playback data including the sound source data and the playback position; and comprising a step of providing the above playback data to the user terminal; Method for providing audio based on captured data.
9. In the audio providing server linked to the user terminal, A database that stores book index data and audio data for books; A communication module for receiving photographed image data of a portion of a book from the user terminal; An analysis module that generates analysis data on text and image features of the above-mentioned captured image data; A search module that matches the above book index data and the above analysis data to generate matching book data; and An audio processing module that determines the playback position of sound data corresponding to the above matching book data and generates playback data. Audio providing server.
10. A database that stores book index data and audio data for books; A shooting module that takes a picture of a part of a book and generates the shooting image data; An analysis module that generates analysis data for at least one of text features and image features of the above-described captured image data; A search module that matches the above book index data and the above analysis data to generate matching book data; and An audio processing module is included that determines the playback position of audio data corresponding to the above matching book data to generate playback data and outputs audio content corresponding to the playback data. User terminal.
Citation Information
Patent Citations
Voice information providing system based on text datatransmission
KR1020050071986A
System for providing contents based on image and methodof the same
KR1020060122690A
The service and method to link web service with book(electronic book)
KR1020120039575A
System and method for providing book-related service based on image
KR1020130137367A
Method, apparatus, and computer program for selecting music based on image
KR1020180122829A