Service providing server that performs voice synthesis based on the emotional elements matched to each paragraph that composes an electronic document and the operating method thereof
The service providing server enhances speech synthesis in electronic documents by matching emotional elements to paragraphs, addressing the challenge of conveying emotions accurately and improving user understanding.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- TJHE HANCOMM INC
- Filing Date
- 2025-01-14
- Publication Date
- 2026-07-21
AI Technical Summary
Existing speech synthesis technologies for electronic documents fail to accurately convey the emotions expressed in different paragraphs, leading to difficulty in understanding the intended emotions and content of the document.
A service providing server that performs speech synthesis based on emotional elements matched to each paragraph, utilizing a model database, matching unit, and generating unit to analyze and synthesize voices corresponding to each paragraph's emotional elements.
Enhances user understanding of the content by providing speech synthesis that accurately reflects the emotions of each paragraph, improving the clarity and emotional depth of the synthesized speech.
Smart Images

Figure PAT00001_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a service providing server that performs speech synthesis based on emotional elements matched to each paragraph constituting an electronic document, and a method of operation thereof. Background Technology
[0002] Recently, with the advancement of speech synthesis technology, a function is being introduced that performs speech synthesis on electronic documents provided by users and provides the resulting output to the users.
[0003] However, the existing speech synthesis function for electronic documents was configured to perform speech synthesis based on a consistent voice for the entire text, regardless of the emotions expressed by each of the multiple paragraphs constituting the document; therefore, there was a problem in that it was difficult for users to understand the emotions intended by each paragraph when listening to the synthesized speech result.
[0004] For example, if one paragraph constituting an electronic document expresses the emotion of 'sadness' and the other expresses the emotion of 'joy,' and if speech synthesis is always performed for both paragraphs based on the same voice regardless of the emotion, there may be a problem in that it is difficult for the user to clearly understand the intent of the person who wrote the electronic document even when listening to the synthesized voice.
[0005] Particularly in the case of literary works such as novels or children's books, given the importance of understanding each emotion, providing speech synthesis results that reflect emotions on a paragraph-by-paragraph basis can be an important factor when performing speech synthesis on electronic documents containing such literary works.
[0006] In this regard, when implementing a function for speech synthesis of electronic documents, if a function is implemented that performs word analysis on each of the multiple paragraphs constituting the electronic document, matches the corresponding sentiment element to each paragraph based on the analysis, and then synthesizes speech for each of the multiple paragraphs using different voices according to the matched sentiment element, it would be of great help to the user in understanding the content of the electronic document while listening to the synthesized speech.
[0007] Therefore, in implementing a function for speech synthesis on electronic documents, research is needed on a technology capable of matching emotional elements to each of the multiple paragraphs constituting the electronic document, performing speech synthesis based on different voices according to the matched emotional elements, and providing the resulting output to the user. The problem to be solved
[0008] The present invention aims to support users utilizing the speech synthesis function for electronic documents in more efficiently understanding the content of the electronic documents by presenting a service providing server that performs speech synthesis based on emotional elements matched to each paragraph constituting an electronic document and a method of operation thereof. means of solving the problem
[0009] A service providing server according to an embodiment of the present invention includes a model database in which data is stored for a voice synthesis model corresponding to each of a plurality of emotional elements—the voice synthesis model corresponding to each of the plurality of emotional elements is a voice synthesis model pre-built based on a voice expressing each emotional element—a matching unit that, when an electronic document containing a plurality of paragraphs is received from a user’s electronic device and a command instructing to perform voice synthesis for the electronic document is received, performs word analysis for each of the plurality of paragraphs inserted in the electronic document and matches one of the plurality of emotional elements corresponding to each paragraph to each of the plurality of paragraphs; a generating unit that generates a synthesized voice corresponding to each of the plurality of paragraphs by loading data for a voice synthesis model corresponding to the emotional element matched to each paragraph from the model database for each of the plurality of paragraphs and performing a process of processing voice synthesis based on the loaded voice synthesis model; and a transmitting unit that generates a voice file in which the synthesized voices for each of the plurality of paragraphs are sequentially connected and then transmits the generated voice file to the electronic device.
[0010] In addition, the method of operation of a service providing server according to an embodiment of the present invention includes the steps of: maintaining a model database in which data is stored for a voice synthesis model corresponding to each of a plurality of emotional elements—the voice synthesis model corresponding to each of the plurality of emotional elements is a voice synthesis model pre-built based on a voice expressing each emotional element—when an electronic document containing a plurality of paragraphs is received from a user’s electronic device and a command instructing to perform voice synthesis for the electronic document is received, performing word analysis for each of the plurality of paragraphs inserted in the electronic document to match one of the plurality of emotional elements corresponding to each paragraph to each of the plurality of paragraphs; for each of the plurality of paragraphs, loading data for a voice synthesis model corresponding to the emotional element matched to each paragraph from the model database and performing a process of processing voice synthesis based on the loaded voice synthesis model to generate a synthesized voice corresponding to each of the plurality of paragraphs; and generating a voice file in which the synthesized voices for each of the plurality of paragraphs are sequentially connected and then transmitting the generated voice file to the electronic device. Effects of the invention
[0011] The present invention provides a service providing server that performs speech synthesis based on emotional elements matched to each paragraph constituting an electronic document and a method of operation thereof, thereby supporting users utilizing the speech synthesis function for electronic documents to understand the content of the electronic document more efficiently. Brief explanation of the drawing
[0012] FIG. 1 is a diagram illustrating the structure of a service providing server according to an embodiment of the present invention. FIG. 2 is a flowchart illustrating the operation method of a service providing server according to an embodiment of the present invention. Specific details for implementing the invention
[0013] Embodiments according to the present invention will be described in detail below with reference to the accompanying drawings. This description is not intended to limit the present invention to specific embodiments and should be understood to include all modifications, equivalents, and substitutions that fall within the spirit and scope of the present invention. Similar reference numerals have been used for similar components in describing each drawing, and unless otherwise defined, all terms used in this specification, including technical or scientific terms, have the same meaning as generally understood by a person skilled in the art to which the present invention pertains.
[0014] In this document, when a part is described as "including" a component, it means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, in various embodiments of the present invention, each component, functional block, or means may be composed of one or more sub-components, and the electrical, electronic, or mechanical functions performed by each component may be implemented by various known devices or mechanical elements, such as electronic circuits, integrated circuits, and ASICs (Application Specific Integrated Circuits), and may be implemented separately or two or more may be integrated into one.
[0015] Meanwhile, the blocks in the attached block diagram or the steps in the flowchart may be interpreted as computer program instructions that perform designated functions by being loaded into the processor or memory of data-processing equipment, such as general-purpose computers, specialized computers, portable notebook computers, and network computers. Since these computer program instructions may be stored in memory provided in a computer device or in memory readable by a computer, the functions described in the blocks in the block diagram or the steps in the flowchart may be produced as manufactured products containing means of instruction to perform them. Furthermore, each block or each step may represent a module, segment, or part of code containing one or more executable instructions for executing a specific logical function(s). Also, it should be noted that in some alternative embodiments, the functions mentioned in the blocks or steps may be executed in a different order than the prescribed order. For example, two blocks or steps shown in succession may be performed substantially simultaneously or in reverse order, and in some cases, some blocks or steps may be omitted.
[0016] FIG. 1 is a diagram illustrating the structure of a service providing server according to an embodiment of the present invention.
[0017] Referring to FIG. 1, the service providing server (110) according to the present invention includes a model database (111), a matching unit (112), a generation unit (113), and a transmission unit (114).
[0018] The model database (111) stores data for speech synthesis models corresponding to each of the multiple emotional elements.
[0019] Here, the aforementioned multiple emotional elements refer to classifications of various emotions that a person can possess, such as 'joy,' 'sadness,' 'anger,' and 'fear.' In this context, the speech synthesis model corresponding to each of the multiple emotional elements refers to a pre-constructed speech synthesis model based on the voice expressing each emotional element. For example, the speech synthesis model corresponding to the emotional element 'joy' refers to a pre-constructed speech synthesis model based on the voice of a person experiencing the emotion corresponding to 'joy,' and the speech synthesis model corresponding to the emotional element 'sadness' refers to a pre-constructed speech synthesis model based on the voice of a person experiencing the emotion corresponding to 'sadness.'
[0020] In this regard, the model database (111) may store data for speech synthesis models corresponding to each of the multiple emotional elements as shown in Table 1 below.
[0022] Multiple emotional elements Speech synthesis model delight Speech Synthesis Model 1 sorrow Speech Synthesis Model 2 anger Speech Synthesis Model 3 ... ...
[0024] When the matching unit (112) receives an electronic document containing multiple paragraphs from the user's electronic device (10) and receives a command instructing to perform speech synthesis for the electronic document, it performs word analysis for each of the multiple paragraphs inserted in the electronic document and matches one of the multiple emotional elements corresponding to each paragraph among the multiple emotional elements for each of the multiple paragraphs.
[0025] In this case, according to one embodiment of the present invention, the matching unit (112) may include a word database (115), a word set database (116), a word set generation unit (117), and an emotion element matching unit (118).
[0026] The word database (115) stores multiple words and word vectors corresponding to each of the multiple words.
[0027] Here, the word vector corresponding to each of the aforementioned multiple words is a pre-embedded vector based on the semantic similarity between words, and is configured such that a higher vector similarity is calculated as the semantic similarity between words increases. For example, if 'Word A' and 'Word B' have similar meanings, word vectors corresponding to 'Word A' and 'Word B' may be assigned so that a high vector similarity is calculated between the two words. Such word vectors may be assigned to each word according to the Word2Vec embedding technique.
[0028] The word set database (116) stores word sets corresponding to each of the aforementioned multiple emotion elements.
[0029] Here, the word set corresponding to each of the aforementioned multiple emotional elements means a set composed of two or more words that are pre-designated as corresponding to each emotional element among the aforementioned multiple words.
[0030] For example, the word set database (116) may have word sets corresponding to each of the multiple emotional elements, such as ‘joy, sadness, anger, ...’, stored as shown in Table 2 below.
[0032] Multiple emotional elements Word set delight {Word X, Word Y, Word Z, ...} sorrow {word a, word b, word c, ...} anger {Word A, Word B, Word C, ...} ... ...
[0034] The word set generation unit (117) generates a word set corresponding to each of the plurality of paragraphs by performing a process of extracting words that match the plurality of words existing in each paragraph for each of the plurality of paragraphs and configuring them into a word set.
[0035] Let us assume that the above multiple paragraphs are composed of ‘Paragraph 1, Paragraph 2, Paragraph 3’. Then, the word set generation unit (117) can generate word sets corresponding to each of ‘Paragraph 1, Paragraph 2, Paragraph 3’ as shown in Table 3 below by performing the process of extracting words that match the above multiple words existing in each paragraph for each of ‘Paragraph 1, Paragraph 2, Paragraph 3’ and configuring them into word sets.
[0037] Multiple paragraphs Word set Paragraph 1 Word 1, Word 2, Word 3, Word 4, Word 5 Paragraph 2 Word 6, Word 7, Word 8, Word 9, Word 10 Paragraph 3 Word 11, Word 12, Word 13, Word 14, Word 15
[0039] The emotion element matching unit (118) calculates the similarity based on word vectors between a word set corresponding to each of the plurality of emotion elements and a word set corresponding to each paragraph for each of the plurality of paragraphs, and then performs a process of matching one of the plurality of emotion elements with the maximum similarity to the emotion element corresponding to the paragraph, thereby matching the emotion element corresponding to each of the plurality of paragraphs.
[0040] At this time, according to one embodiment of the present invention, the emotion element matching unit (118) performs a process of sequentially matching emotion elements for each of the plurality of paragraphs, and when it is the turn to perform the process of matching emotion elements for a first paragraph, which is one of the plurality of paragraphs, the word vectors of words constituting a word set corresponding to each of the plurality of emotion elements and the word vectors of words constituting a word set corresponding to the first paragraph are confirmed by referring to the word database (115), and then for each of the plurality of emotion elements, all vector similarities between each of the words constituting the word set corresponding to each emotion element and each of the words constituting the word set corresponding to the first paragraph are calculated, and when the calculated vector similarities are sorted in ascending order and the quartiles are calculated, the overall average of the remaining vector similarities, excluding the vector similarities corresponding to the first quartile or lower, is calculated as the similarity corresponding to each emotion element, thereby calculating the similarity between the word set corresponding to each of the plurality of emotion elements and the word set corresponding to the first paragraph, and then the plurality Among the emotional elements, any one emotional element with the maximum similarity can be matched to the emotional element corresponding to the first paragraph above.
[0041] Assuming that, as in the example described above, the plurality of paragraphs are composed of ‘Paragraph 1, Paragraph 2, Paragraph 3’, and that a word set corresponding to each of the plurality of paragraphs is generated by the word set generation unit (117) as shown in Table 3, and that the plurality of emotional elements are called ‘Joy, Sadness, Anger, ...’, and that a word set corresponding to each of the plurality of emotional elements is stored as shown in Table 2, the process of the emotional element matching unit (118) matching the emotional elements corresponding to each of the plurality of paragraphs is explained as follows.
[0042] First, the emotion element matching unit (118) can perform the process of sequentially matching emotion elements to each of the plurality of paragraphs, 'paragraph 1, paragraph 2, paragraph 3'.
[0043] At this time, when it is said that the time has come to perform the process of matching emotional elements to 'Paragraph 1', which is one of the plurality of paragraphs above, the emotional element matching unit (118) can refer to the word database (115) to check the word vectors of the words constituting the word set corresponding to each of the plurality of emotional elements, such as checking the word vectors of 'word X, word Y, word Z, ...' which constitute the word set corresponding to 'joy', 'word a, word b, word c, ...' which constitute the word set corresponding to 'sadness', and 'word a, word b, word c, ...' which constitute the word set corresponding to 'anger'.
[0044] In addition, the emotion element matching unit (118) can also check word vectors for ‘word 1, word 2, word 3, word 4, word 5’, which are words constituting a word set corresponding to ‘paragraph 1’, by referring to the word database (115).
[0045] Then, the emotion element matching unit (118) can perform a process of calculating all vector similarities between each word constituting the word set corresponding to each emotion element and each word constituting the word set corresponding to 'paragraph 1' for each of the plurality of emotion elements, and when the calculated vector similarities are sorted in ascending order to calculate the quartiles, the overall average of the remaining vector similarities, excluding the vector similarities corresponding to the first quartile or lower, can be calculated as the similarity corresponding to each emotion element.
[0046] First, the emotion element matching unit (118) can calculate all vector similarities between each of the words that make up the word set corresponding to 'joy' among the plurality of emotion elements, namely 'word X, word Y, word Z, ...', and each of the words that make up the word set corresponding to 'paragraph 1', namely 'word 1, word 2, word 3, word 4, word 5'.
[0047] Specifically, the emotion element matching part (118) is vector similarity between 'word X' and 'word 1', vector similarity between 'word Y' and 'word 1', vector similarity between 'word Z' and 'word 1', vector similarity between 'word X' and 'word 2', vector similarity between 'word Y' and 'word 2', vector similarity between 'word Z' and 'word 2', vector similarity between 'word X' and 'word 3', vector similarity between 'word Y' and 'word 3', vector similarity between 'word Z' and 'word 3', vector similarity between 'word X' and 'word 4', vector similarity between 'word Y' and 'word 4', vector similarity between 'word Z' and 'word 4', vector similarity between 'word X' and 'word 5', and vector similarity between 'word Y' and 'word 5'. Similarities, vector similarities between 'word Z' and 'word 5', ..., vector similarities between each word in the word set corresponding to 'joy' and each word in the word set corresponding to 'paragraph 1' can all be calculated.
[0048] After that, the emotion element matching unit (118) can calculate the quartiles by sorting the calculated vector similarities in ascending order.
[0049] Here, quartiles refer to data sorted in ascending order divided into four equal parts, where the first quartile is the variable value corresponding to the 1 / 4 place and the third quartile is the variable value corresponding to the 3 / 4 place.
[0050] At this time, the vector similarities calculated between each of the words constituting the word set corresponding to 'joy' and each of the words constituting the word set corresponding to 'paragraph 1' are called 'V1, V2, V3, V4, V5, V6, V7, V8, V9, V10, V11, V12, V13, V14, V15, ...', and the vector similarities corresponding to the first quartile or lower among these are called 'V4, V7'. Then, the emotion element matching unit (118) can calculate the overall average of the remaining vector similarities, 'V1, V2, V3, V5, V6, V8, V9, V10, V11, V12, V13, V14, V15, ...', excluding the vector similarities corresponding to the first quartile or lower among the vector similarities, as the similarity corresponding to 'joy'.
[0051] Additionally, the emotion element matching unit (118) can calculate all vector similarities between each of the words that make up the word set corresponding to 'sadness' among the plurality of emotion elements, namely 'word a, word b, word c, ...', and each of the words that make up the word set corresponding to 'paragraph 1', namely 'word 1, word 2, word 3, word 4, word 5'.
[0052] Specifically, the emotion element matching part (118) is vector similarity between 'word a' and 'word 1', vector similarity between 'word b' and 'word 1', vector similarity between 'word c' and 'word 1', vector similarity between 'word a' and 'word 2', vector similarity between 'word b' and 'word 2', vector similarity between 'word c' and 'word 2', vector similarity between 'word a' and 'word 3', vector similarity between 'word b' and 'word 3', vector similarity between 'word c' and 'word 3', vector similarity between 'word a' and 'word 4', vector similarity between 'word b' and 'word 4', vector similarity between 'word c' and 'word 4', vector similarity between 'word a' and 'word 5', and vector similarity between 'word b' and 'word 5'. Vector similarities between each word in the word set corresponding to 'sadness' and each word in the word set corresponding to 'paragraph 1' can all be calculated, such as similarity, vector similarity between 'word c' and 'word 5', ....
[0053] After that, the emotion element matching unit (118) can calculate the quartiles by sorting the calculated vector similarities in ascending order.
[0054] At this time, the vector similarities calculated between each of the words constituting the word set corresponding to 'sadness' and each of the words constituting the word set corresponding to 'paragraph 1' are referred to as 'V16, V17, V18, V19, V20, V21, V22, V23, V24, V25, V26, V27, V28, V29, V30, ...', and the vector similarities corresponding to the first quartile or lower among these are referred to as 'V16, V18'. Then, the emotion element matching unit (118) can calculate the overall average of the remaining vector similarities, 'V17, V19, V20, V21, V22, V23, V24, V25, V26, V27, V28, V29, V30, ...', excluding the vector similarities corresponding to the first quartile or lower among the vector similarities, as the similarity corresponding to 'sadness'.
[0055] Additionally, the emotion element matching unit (118) can calculate all vector similarities between each of the words that make up the word set corresponding to 'anger' among the plurality of emotion elements, namely 'word a, word b, word c, ...', and each of the words that make up the word set corresponding to 'paragraph 1', namely 'word 1, word 2, word 3, word 4, word 5'.
[0056] Specifically, the emotion element matching part (118) has a vector similarity between 'word a' and 'word 1', a vector similarity between 'word b' and 'word 1', a vector similarity between 'word c' and 'word 1', a vector similarity between 'word a' and 'word 2', a vector similarity between 'word b' and 'word 2', a vector similarity between 'word c' and 'word 2', a vector similarity between 'word a' and 'word 3', a vector similarity between 'word b' and 'word 3', a vector similarity between 'word c' and 'word 3', a vector similarity between 'word a' and 'word 4', a vector similarity between 'word b' and 'word 4', a vector similarity between 'word c' and 'word 4', a vector similarity between 'word a' and 'word 5', and a vector similarity between 'word b' and 'word 5'. Similarities, vector similarities between 'word 5' and 'word 5', ..., vector similarities between each word in the word set corresponding to 'anger' and each word in the word set corresponding to 'paragraph 1' can all be calculated.
[0057] After that, the emotion element matching unit (118) can calculate the quartiles by sorting the calculated vector similarities in ascending order.
[0058] At this time, the vector similarities calculated between each word constituting the word set corresponding to 'anger' and each word constituting the word set corresponding to 'paragraph 1' are referred to as 'V31, V32, V33, V34, V35, V36, V37, V38, V39, V40, V41, V42, V43, V44, V45, ...', and the vector similarities corresponding to the first quartile or lower among these are referred to as 'V31, V45'. Then, the emotion element matching unit (118) can calculate the overall average of the remaining vector similarities, 'V32, V33, V34, V35, V36, V37, V38, V39, V40, V41, V42, V43, V44, ...', excluding the vector similarities corresponding to the first quartile or lower among the vector similarities, as the similarity corresponding to 'anger'.
[0059] In this way, the emotion element matching unit (118) can calculate the similarity corresponding to each of the plurality of emotion elements, 'joy, sadness, anger, ...'.
[0060] When the calculation of similarity corresponding to each of the aforementioned multiple emotional elements, ‘joy, sadness, anger, ...’ is completed, the emotional element matching unit (118) can match one of the multiple emotional elements with the maximum similarity to the emotional element corresponding to ‘paragraph 1’.
[0061] In this regard, among the aforementioned multiple emotional elements, ‘joy, sadness, anger, ...’, if one emotional element with the maximum similarity is called ‘sadness’, the emotional element matching unit (118) can match ‘sadness’ to the emotional element corresponding to ‘paragraph 1’.
[0062] When the process of matching emotional elements corresponding to 'Paragraph 1' is completed and it is time to perform the process of matching emotional elements to 'Paragraph 2', which is one of the plurality of paragraphs, the emotional element matching unit (118) can check the word vectors of the words constituting the word set corresponding to each of the plurality of emotional elements by referring to the word database (115), such as checking the word vectors of the words constituting the word set corresponding to 'joy', such as 'word X, word Y, word Z, ...', the words constituting the word set corresponding to 'sadness', such as 'word a, word b, word c, ...', and the words constituting the word set corresponding to 'anger', such as 'word a, word b, word c, ...'.
[0063] In addition, the emotion element matching unit (118) can also check word vectors for ‘word 6, word 7, word 8, word 9, word 10’, which are words constituting the word set corresponding to ‘paragraph 2’, by referring to the word database (115).
[0064] Then, the emotion element matching unit (118) can perform a process of calculating all vector similarities between each word constituting the word set corresponding to each emotion element and each word constituting the word set corresponding to 'paragraph 2' for each of the plurality of emotion elements, and when the calculated vector similarities are sorted in ascending order to calculate the quartiles, the overall average of the remaining vector similarities, excluding the vector similarities corresponding to the first quartile or lower, can be calculated as the similarity corresponding to each emotion element.
[0065] First, the emotion element matching unit (118) can calculate all vector similarities between each of the words that make up the word set corresponding to 'joy' among the plurality of emotion elements, namely 'word X, word Y, word Z, ...', and each of the words that make up the word set corresponding to 'paragraph 2', namely 'word 6, word 7, word 8, word 9, word 10'.
[0066] Specifically, the emotion element matching section (118) includes vector similarity between 'word X' and 'word 6', vector similarity between 'word Y' and 'word 6', vector similarity between 'word Z' and 'word 6', vector similarity between 'word X' and 'word 7', vector similarity between 'word Y' and 'word 7', vector similarity between 'word Z' and 'word 7', vector similarity between 'word X' and 'word 8', vector similarity between 'word Y' and 'word 8', vector similarity between 'word Z' and 'word 8', vector similarity between 'word X' and 'word 9', vector similarity between 'word Y' and 'word 9', vector similarity between 'word Z' and 'word 9', vector similarity between 'word X' and 'word 10', and vector similarity between 'word Y' and 'word 10' Vector similarity between, vector similarity between 'word Z' and 'word 10', ..., and vector similarity between each word in the word set corresponding to 'joy' and each word in the word set corresponding to 'paragraph 2' can all be calculated.
[0067] After that, the emotion element matching unit (118) can calculate the quartiles by sorting the calculated vector similarities in ascending order.
[0068] At this time, the vector similarities calculated between each of the words constituting the word set corresponding to 'joy' among the plurality of emotional elements and each of the words constituting the word set corresponding to 'paragraph 2' are called 'C1, C2, C3, C4, C5, C6, C7, C8, C9, C10, C11, C12, C13, C14, C15, ...', and the vector similarities corresponding to the first quartile or lower among these are called 'C5, C8', then the emotional element matching unit (118) can calculate the overall average of the remaining vector similarities, 'C1, C2, C3, C4, C6, C7, C9, C10, C11, C12, C13, C14, C15, ...', excluding the vector similarities corresponding to the first quartile or lower among the vector similarities, as the similarity corresponding to 'joy'.
[0069] Additionally, the emotion element matching unit (118) can calculate all vector similarities between each of the words that make up the word set corresponding to 'sadness' among the plurality of emotion elements, namely 'word a, word b, word c, ...', and each of the words that make up the word set corresponding to 'paragraph 2', namely 'word 6, word 7, word 8, word 9, word 10'.
[0070] Specifically, the emotion element matching section (118) includes vector similarity between 'word a' and 'word 6', vector similarity between 'word b' and 'word 6', vector similarity between 'word c' and 'word 6', vector similarity between 'word a' and 'word 7', vector similarity between 'word b' and 'word 7', vector similarity between 'word c' and 'word 7', vector similarity between 'word a' and 'word 8', vector similarity between 'word b' and 'word 8', vector similarity between 'word c' and 'word 8', vector similarity between 'word a' and 'word 9', vector similarity between 'word b' and 'word 9', vector similarity between 'word c' and 'word 9', vector similarity between 'word a' and 'word 10', and vector similarity between 'word b' and 'word 10' Vector similarity between, vector similarity between 'word c' and 'word 10', ..., and vector similarity between each word in the word set corresponding to 'sadness' and each word in the word set corresponding to 'paragraph 2' can all be calculated.
[0071] After that, the emotion element matching unit (118) can calculate the quartiles by sorting the calculated vector similarities in ascending order.
[0072] At this time, when the vector similarities calculated between each of the words constituting the word set corresponding to 'sadness' among the plurality of emotional elements and each of the words constituting the word set corresponding to 'paragraph 2' are denoted as 'C16, C17, C18, C19, C20, C21, C22, C23, C24, C25, C26, C27, C28, C29, C30, ...' and the vector similarities corresponding to the first quartile or lower among these are denoted as 'C17, C19', the emotional element matching unit (118) can calculate the overall average of the remaining vector similarities, 'C16, C18, C20, C21, C22, C23, C24, C25, C26, C27, C28, C29, C30, ...', excluding the vector similarities corresponding to the first quartile or lower among the vector similarities, as the similarity corresponding to 'sadness'. there is.
[0073] Additionally, the emotion element matching unit (118) can calculate all vector similarities between each of the words that make up the word set corresponding to 'anger' among the plurality of emotion elements, namely 'word a, word b, word c, ...', and each of the words that make up the word set corresponding to 'paragraph 2', namely 'word 6, word 7, word 8, word 9, word 10'.
[0074] Specifically, the emotion element matching section (118) includes vector similarity between 'word a' and 'word 6', vector similarity between 'word b' and 'word 6', vector similarity between 'word c' and 'word 6', vector similarity between 'word a' and 'word 7', vector similarity between 'word b' and 'word 7', vector similarity between 'word c' and 'word 7', vector similarity between 'word a' and 'word 8', vector similarity between 'word b' and 'word 8', vector similarity between 'word c' and 'word 8', vector similarity between 'word a' and 'word 9', vector similarity between 'word b' and 'word 9', vector similarity between 'word c' and 'word 9', vector similarity between 'word a' and 'word 10', and vector similarity between 'word b' and 'word 10' Vector similarity between, vector similarity between 'word 10' and 'word 10', ..., and vector similarity between each word in the word set corresponding to 'anger' and each word in the word set corresponding to 'paragraph 2' can all be calculated.
[0075] After that, the emotion element matching unit (118) can calculate the quartiles by sorting the calculated vector similarities in ascending order.
[0076] At this time, when the vector similarities calculated between each of the words constituting the word set corresponding to 'anger' among the plurality of emotion elements and each of the words constituting the word set corresponding to 'paragraph 2' are denoted as 'C31, C32, C33, C34, C35, C36, C37, C38, C39, C40, C41, C42, C43, C44, C45, ...' and the vector similarities corresponding to the first quartile or lower among these are denoted as 'C32, C45', the emotion element matching unit (118) can calculate the overall average of the remaining vector similarities, 'C31, C33, C34, C35, C36, C37, C38, C39, C40, C41, C42, C43, C44, ...', excluding the vector similarities corresponding to the first quartile or lower among the vector similarities, as the similarity corresponding to 'anger'. there is.
[0077] In this way, the emotion element matching unit (118) can calculate the similarity corresponding to each of the plurality of emotion elements, 'joy, sadness, anger, ...'.
[0078] When the calculation of similarity corresponding to each of the aforementioned multiple emotional elements, ‘joy, sadness, anger, ...’ is completed, the emotional element matching unit (118) can match one of the multiple emotional elements with the maximum similarity to the emotional element corresponding to ‘paragraph 2’.
[0079] In this regard, among the above multiple emotional elements, ‘joy, sadness, anger, ...’, if one emotional element with the maximum similarity is called ‘anger’, the emotional element matching unit (118) can match ‘anger’ to the emotional element corresponding to ‘paragraph 2’.
[0080] When the process of matching emotional elements corresponding to 'Paragraph 2' is completed and it is time to perform the process of matching emotional elements to 'Paragraph 3', which is one of the plurality of paragraphs, the emotional element matching unit (118) can check the word vectors of the words constituting the word set corresponding to each of the plurality of emotional elements by referring to the word database (115), such as checking the word vectors of the words constituting the word set corresponding to 'joy', such as 'word X, word Y, word Z, ...', the words constituting the word set corresponding to 'sadness', such as 'word a, word b, word c, ...', and the words constituting the word set corresponding to 'anger', such as 'word a, word b, word c, ...'.
[0081] In addition, the emotion element matching unit (118) can also check word vectors for ‘word 11, word 12, word 13, word 14, word 15’, which are words constituting the word set corresponding to ‘paragraph 3’, by referring to the word database (115).
[0082] Then, the emotion element matching unit (118) can perform a process of calculating all vector similarities between each word constituting the word set corresponding to each emotion element and each word constituting the word set corresponding to 'paragraph 3' for each of the plurality of emotion elements, and when the calculated vector similarities are sorted in ascending order to calculate the quartiles, the overall average of the remaining vector similarities, excluding the vector similarities corresponding to the first quartile or lower, can be calculated as the similarity corresponding to each emotion element.
[0083] First, the emotion element matching unit (118) can calculate all vector similarities between each of the words that make up the word set corresponding to 'joy' among the plurality of emotion elements, namely 'word X, word Y, word Z, ...', and each of the words that make up the word set corresponding to 'paragraph 3', namely 'word 11, word 12, word 13, word 14, word 15'.
[0084] Specifically, the emotion element matching section (118) includes vector similarity between 'word X' and 'word 11', vector similarity between 'word Y' and 'word 11', vector similarity between 'word Z' and 'word 11', vector similarity between 'word X' and 'word 12', vector similarity between 'word Y' and 'word 12', vector similarity between 'word Z' and 'word 12', vector similarity between 'word X' and 'word 13', vector similarity between 'word Y' and 'word 13', vector similarity between 'word Z' and 'word 13', vector similarity between 'word X' and 'word 14', vector similarity between 'word Y' and 'word 14', vector similarity between 'word Z' and 'word 14', vector similarity between 'word X' and 'word 15', Vector similarity between 'word Y' and 'word 15', vector similarity between 'word Z' and 'word 15', ..., and vector similarity between each word in the word set corresponding to 'joy' and each word in the word set corresponding to 'paragraph 3' can all be calculated.
[0085] After that, the emotion element matching unit (118) can calculate the quartiles by sorting the calculated vector similarities in ascending order.
[0086] At this time, the vector similarities calculated between each word constituting the word set corresponding to 'joy' among the plurality of emotional elements and each word constituting the word set corresponding to 'paragraph 3' are denoted as 'T1, T2, T3, T4, T5, T6, T7, T8, T9, T10, T11, T12, T13, T14, T15, ...', and the vector similarities corresponding to the first quartile or lower among these are denoted as 'T3, T6'. Then, the emotional element matching unit (118) can calculate the overall average of the remaining vector similarities, 'T1, T2, T4, T5, T7, T8, T9, T10, T11, T12, T13, T14, T15, ...', excluding the vector similarities corresponding to the first quartile or lower among the vector similarities, as the similarity corresponding to 'joy'.
[0087] Additionally, the emotion element matching unit (118) can calculate all vector similarities between each of the words 'word a, word b, word c, ...' that constitute the word set corresponding to 'sadness' among the plurality of emotion elements, and each of the words 'word 11, word 12, word 13, word 14, word 15' that constitute the word set corresponding to 'paragraph 3'.
[0088] Specifically, the emotion element matching part (118) includes vector similarity between 'word a' and 'word 11', vector similarity between 'word b' and 'word 11', vector similarity between 'word c' and 'word 11', vector similarity between 'word a' and 'word 12', vector similarity between 'word b' and 'word 12', vector similarity between 'word c' and 'word 12', vector similarity between 'word a' and 'word 13', vector similarity between 'word b' and 'word 13', vector similarity between 'word c' and 'word 13', vector similarity between 'word a' and 'word 14', vector similarity between 'word b' and 'word 14', vector similarity between 'word c' and 'word 14', and vector similarity between 'word a' and 'word 15'. Vector similarity between 'word b' and 'word 15', vector similarity between 'word c' and 'word 15', ..., and vector similarity between each word in the word set corresponding to 'sadness' and each word in the word set corresponding to 'paragraph 3' can all be calculated.
[0089] After that, the emotion element matching unit (118) can calculate the quartiles by sorting the calculated vector similarities in ascending order.
[0090] At this time, when the vector similarities calculated between each of the words constituting the word set corresponding to 'sadness' among the plurality of emotion elements and each of the words constituting the word set corresponding to 'paragraph 3' are denoted as 'T16, T17, T18, T19, T20, T21, T22, T23, T24, T25, T26, T27, T28, T29, T30, ...' and the vector similarities corresponding to the first quartile or lower among these are denoted as 'T16, T17', the emotion element matching unit (118) can calculate the overall average of the remaining vector similarities, 'T18, T19, T20, T21, T22, T23, T24, T25, T26, T27, T28, T29, T30, ...', excluding the vector similarities corresponding to the first quartile or lower among the vector similarities, as the similarity corresponding to 'sadness'. there is.
[0091] Additionally, the emotion element matching unit (118) can calculate all vector similarities between each of the words that make up the word set corresponding to 'anger' among the plurality of emotion elements, namely 'word a, word b, word c, ...', and each of the words that make up the word set corresponding to 'paragraph 3', namely 'word 11, word 12, word 13, word 14, word 15'.
[0092] Specifically, the emotion element matching part (118) includes vector similarity between 'word a' and 'word 11', vector similarity between 'word b' and 'word 11', vector similarity between 'word c' and 'word 11', vector similarity between 'word a' and 'word 12', vector similarity between 'word b' and 'word 12', vector similarity between 'word c' and 'word 12', vector similarity between 'word a' and 'word 13', vector similarity between 'word b' and 'word 13', vector similarity between 'word c' and 'word 13', vector similarity between 'word a' and 'word 14', vector similarity between 'word b' and 'word 14', vector similarity between 'word c' and 'word 14', vector similarity between 'word a' and 'word 15', Vector similarity between 'word Na' and 'word 15', vector similarity between 'word Da' and 'word 15', ..., and vector similarity between each word in the word set corresponding to 'anger' and each word in the word set corresponding to 'paragraph 3' can all be calculated.
[0093] After that, the emotion element matching unit (118) can calculate the quartiles by sorting the calculated vector similarities in ascending order.
[0094] At this time, when the vector similarities calculated between each word constituting the word set corresponding to 'anger' among the plurality of emotion elements and each word constituting the word set corresponding to 'paragraph 3' are denoted as 'T31, T32, T33, T34, T35, T36, T37, T38, T39, T40, T41, T42, T43, T44, T45, ...' and the vector similarities corresponding to the first quartile or lower among these are denoted as 'T31, T44', the emotion element matching unit (118) can calculate the overall average of the remaining vector similarities, excluding the vector similarities corresponding to the first quartile or lower among the vector similarities, namely 'T32, T33, T34, T35, T36, T37, T38, T39, T40, T41, T42, T43, T45, ...', as the similarity corresponding to 'anger'. there is.
[0095] In this way, the emotion element matching unit (118) can calculate the similarity corresponding to each of the plurality of emotion elements, 'joy, sadness, anger, ...'.
[0096] When the calculation of similarity corresponding to each of the aforementioned multiple emotional elements, 'joy, sadness, anger, ...' is completed, the emotional element matching unit (118) can match one of the multiple emotional elements with the maximum similarity to the emotional element corresponding to 'paragraph 3'.
[0097] In this regard, among the above multiple emotional elements, ‘joy, sadness, anger, ...’, if one emotional element with the maximum similarity is called ‘joy’, the emotional element matching unit (118) can match ‘joy’ to the emotional element corresponding to ‘paragraph 3’.
[0098] In this way, the emotion element matching unit (118) performs a process of matching one emotion element with the maximum similarity among the plurality of emotion elements, ‘joy, sadness, anger, ...’, to each of the plurality of paragraphs, ‘paragraph 1, paragraph 2, paragraph 3’, as the emotion element corresponding to the paragraph, thereby being able to match ‘sadness’ to ‘paragraph 1’, match ‘anger’ to ‘paragraph 2’, and match ‘joy’ to ‘paragraph 3’.
[0099] The generation unit (113) generates a synthesized voice corresponding to each of the plurality of paragraphs by performing a process of processing voice synthesis based on the loaded voice synthesis model, by loading data for a voice synthesis model corresponding to an emotion element matched to each paragraph from a model database (111) for each of the plurality of paragraphs.
[0100] Let us assume that the above multiple paragraphs are composed of ‘Paragraph 1, Paragraph 2, Paragraph 3’, that the emotional element matched to ‘Paragraph 1’ is ‘sadness’, that the emotional element matched to ‘Paragraph 2’ is ‘anger’, and that the emotional element matched to ‘Paragraph 3’ is ‘joy’, and that data for a speech synthesis model corresponding to each of the multiple emotional elements is stored in the model database (111) as shown in Table 1 above.
[0101] Then, the generation unit (113) can generate ‘synthetic voice 1’ corresponding to ‘paragraph 1’ by loading data from the model database (111) for ‘paragraph 1’ among the plurality of paragraphs, ‘paragraph 1, paragraph 2, paragraph 3’, and processing voice synthesis based on the loaded ‘speech synthesis model 2’, which is a voice synthesis model corresponding to ‘paragraph 1’.
[0102] Additionally, the generation unit (113) can generate a synthesized voice corresponding to 'paragraph 2' by loading data for 'speech synthesis model 3', which is a speech synthesis model corresponding to 'anger', an emotion element matched to 'paragraph 2', from the model database (111) for 'paragraph 2' among the plurality of paragraphs 'paragraph 1, paragraph 2, paragraph 3', and by performing a process of processing speech synthesis based on the loaded 'speech synthesis model 3'.
[0103] Additionally, the generation unit (113) can generate a synthesized voice corresponding to 'paragraph 3' by loading data for 'speech synthesis model 1', which is a speech synthesis model corresponding to 'joy', an emotion element matched to 'paragraph 3', from the model database (111) for 'paragraph 3' among the plurality of paragraphs 'paragraph 1, paragraph 2, paragraph 3', and by performing a process of processing speech synthesis based on the loaded 'speech synthesis model 1'.
[0104] In this way, the generating unit (113) can generate a synthetic voice corresponding to 'paragraph 1' for each of the plurality of paragraphs, 'paragraph 1, paragraph 2, paragraph 3', a synthetic voice corresponding to 'paragraph 1', a synthetic voice corresponding to 'paragraph 2', and a synthetic voice corresponding to 'paragraph 3'.
[0105] The transmission unit (114) generates a voice file in which the synthesized voices for each of the plurality of paragraphs are sequentially connected, and then transmits the generated voice file to the electronic device (10).
[0106] At this time, according to one embodiment of the present invention, when the transmission unit (114) generates a voice file in which synthesized voices for each of the plurality of paragraphs are sequentially connected, in the process of sequentially connecting the synthesized voices corresponding to each of the plurality of paragraphs, a silent interval of a preset time is included between the synthesized voices corresponding to each paragraph, thereby generating a voice file that includes silent intervals between the synthesized voices corresponding to each paragraph.
[0107] Let us assume that the above multiple paragraphs are composed of ‘Paragraph 1, Paragraph 2, Paragraph 3’, that the synthesized voice for ‘Paragraph 1’ is generated as ‘Synthetic Voice 1’ by the generation unit (113), that the synthesized voice for ‘Paragraph 2’ is generated as ‘Synthetic Voice 2’, that the synthesized voice for ‘Paragraph 3’ is generated as ‘Synthetic Voice 3’, and that the above preset time is ‘2 seconds’.
[0108] Then, when the transmission unit (114) generates a voice file in which the synthesized voices for each of the plurality of paragraphs, 'paragraph 1, paragraph 2, paragraph 3', 'synthetic voice 1, synthetic voice 2, synthetic voice 3', are sequentially connected, in the process of sequentially connecting 'synthetic voice 1, synthetic voice 2, synthetic voice 3', a silent interval of '2 seconds', which is a preset time, is included between 'synthetic voice 1' and 'synthetic voice 2' and between 'synthetic voice 2' and 'synthetic voice 3', thereby generating a voice file that includes a silent interval between the synthesized voices corresponding to each paragraph, 'synthetic voice 1, synthetic voice 2, synthetic voice 3', and then transmits the generated voice file to the electronic device (10).
[0109] In this case, according to one embodiment of the present invention, the service providing server (110) may further include a summary generation unit (119) and a table transmission unit (120).
[0110] After a voice file is transmitted to the electronic device (10) through the transmission unit (114), when the summary generation unit (119) receives a request command from the electronic device (10) to provide an emotional element analysis table that analyzes and records emotional elements for each of the plurality of paragraphs, it generates a summary by inputting the text contained in each of the plurality of paragraphs as input to a preset summary generation model.
[0111] Here, the aforementioned summary generation model refers to a generative artificial intelligence model that takes text as input and automatically generates a summary corresponding to the input text based on a pre-built Large Language Model (LM).
[0112] In this regard, let's assume that the above multiple paragraphs consist of 'Paragraph 1, Paragraph 2, and Paragraph 3', let the text included in 'Paragraph 1' be called 'Text 1', the text included in 'Paragraph 2' be called 'Text 2', and the text included in 'Paragraph 3' be called 'Text 3'.
[0113] Then, when the summary generation unit (119) receives a request command to provide an emotional element analysis table that analyzes and records emotional elements for each of the plurality of paragraphs, 'paragraph 1, paragraph 2, paragraph 3', from the electronic device (10) after the voice file is transmitted to the electronic device (10) through the transmission unit (114), the summary generation unit (119) can generate 'summary 1' by inputting 'text 1', which is text included in 'paragraph 1', to the summary generation model, and can generate 'summary 2' by inputting 'text 2', which is text included in 'paragraph 2', to the summary generation model, and can generate 'summary 3' by inputting 'text 3', which is text included in 'paragraph 3', to the summary generation model.
[0114] The table transmission unit (120) generates an emotional element analysis table in which a summary corresponding to each of the plurality of paragraphs and an emotional element corresponding to each paragraph are recorded in correspondence with each other, and then transmits the generated emotional element analysis table to the electronic device (10).
[0115] Let us assume that, as in the example described above, the above multiple paragraphs are composed of ‘Paragraph 1, Paragraph 2, Paragraph 3’, and that a summary ‘Summary 1’ corresponding to ‘Paragraph 1’ is generated by the summary generation unit (119), a summary ‘Summary 2’ corresponding to ‘Paragraph 2’ is generated, and a summary ‘Summary 3’ corresponding to ‘Paragraph 3’ is generated, and that an emotional element corresponding to ‘Paragraph 1’ is matched to ‘sadness’ by the matching unit (112), an emotional element corresponding to ‘Paragraph 2’ is matched to ‘anger’, and an emotional element corresponding to ‘Paragraph 3’ is matched to ‘joy’.
[0116] Then, the table transmission unit (120) can generate an emotional element analysis table as shown in Table 4 below, in which a summary corresponding to each of the plurality of paragraphs, 'paragraph 1, paragraph 2, paragraph 3', and an emotional element corresponding to each paragraph are recorded in correspondence with each other, and then transmit the generated emotional element analysis table to the electronic device (10).
[0118] Paragraph number Summary Emotional elements 1 Summary 1 sorrow 2 Summary 2 anger 3 Summary 3 delight
[0120] Through this, the user can grasp at a glance the summary of each of the multiple paragraphs constituting the electronic document, namely 'Paragraph 1, Paragraph 2, and Paragraph 3', and the emotional elements corresponding to each paragraph.
[0121] FIG. 2 is a flowchart illustrating the operation method of a service providing server according to an embodiment of the present invention.
[0122] In step (S210), a model database is maintained in which data for a speech synthesis model corresponding to each of a plurality of emotional elements (the speech synthesis model corresponding to each of the plurality of emotional elements is a pre-built speech synthesis model based on a voice expressing each emotional element) is stored.
[0123] In step (S220), when an electronic document containing multiple paragraphs is received from a user's electronic device and a command is received instructing to perform speech synthesis on the electronic document, word analysis is performed on each of the multiple paragraphs inserted in the electronic document, and for each of the multiple paragraphs, one of the multiple emotional elements corresponding to each paragraph is matched.
[0124] In step (S230), for each of the plurality of paragraphs, data for a speech synthesis model corresponding to an emotion element matched to each paragraph is loaded from the model database, and a process of processing speech synthesis based on the loaded speech synthesis model is performed to generate a synthesized speech corresponding to each of the plurality of paragraphs.
[0125] In step (S240), a voice file is generated in which the synthesized voices for each of the plurality of paragraphs are sequentially connected, and then the generated voice file is transmitted to the electronic device.
[0126] At this time, according to one embodiment of the present invention, step (S220) comprises: maintaining a word database in which a plurality of words and a word vector corresponding to each of the plurality of words are stored (the word vector corresponding to each of the plurality of words is a pre-embedded vector based on semantic similarity between words, and is specified such that a higher vector similarity is calculated as the semantic similarity between words is higher); maintaining a word set database in which a word set corresponding to each of the plurality of emotion elements is stored (the word set corresponding to each of the plurality of emotion elements is a set composed of two or more words pre-specified as corresponding to each emotion element among the plurality of words); for each of the plurality of paragraphs, a process of extracting words that match the plurality of words existing in each paragraph and configuring them into a word set is performed to generate a word set corresponding to each of the plurality of paragraphs; and for each of the plurality of paragraphs, after calculating the similarity based on the word vector between the word set corresponding to each of the plurality of emotion elements and the word set corresponding to each paragraph, one of the emotion elements with the maximum similarity among the plurality of emotion elements is applied to the paragraph. By performing a process of matching with corresponding emotional elements, the method may include a step of matching emotional elements corresponding to each of the aforementioned plurality of paragraphs.
[0127] In this case, according to one embodiment of the present invention, the step of matching corresponding emotional elements is to sequentially perform the process of matching emotional elements for each of the plurality of paragraphs, wherein when it is the turn to perform the process of matching emotional elements for a first paragraph, which is one of the plurality of paragraphs, the word vectors of the words constituting the word set corresponding to each of the plurality of emotional elements and the word vectors of the words constituting the word set corresponding to the first paragraph are confirmed by referring to the word database, and then for each of the plurality of emotional elements, all vector similarities between each of the words constituting the word set corresponding to each emotional element and each of the words constituting the word set corresponding to the first paragraph are calculated, and when the calculated vector similarities are sorted in ascending order and quartiles are calculated, the overall average of the remaining vector similarities, excluding the vector similarities corresponding to the first quartile or lower, is calculated as the similarity corresponding to each emotional element, thereby performing the process of calculating the similarity between the word set corresponding to each of the plurality of emotional elements and the word set corresponding to the first paragraph, and among the plurality of emotional elements, the one with the maximum similarity Any one emotional element can be matched to the emotional element corresponding to the first paragraph above.
[0128] In addition, according to one embodiment of the present invention, in step (S240), when generating a voice file in which synthesized voices for each of the plurality of paragraphs are sequentially connected, a silent interval of a preset time is included between the synthesized voices corresponding to each paragraph during the process of sequentially connecting the synthesized voices corresponding to each of the plurality of paragraphs, thereby generating a voice file that includes silent intervals between the synthesized voices corresponding to each paragraph.
[0129] In addition, according to one embodiment of the present invention, the method of operation of the service providing server may further include the step of, after a voice file is transmitted to the electronic device through step (S240), when a request command to provide an emotional element analysis table is received from the electronic device, which analyzes and records emotional elements for each of the plurality of paragraphs, inputting the text contained in each of the plurality of paragraphs as input to a pre-set summary generation model (the summary generation model is a generative artificial intelligence model that receives text input and automatically generates a summary corresponding to the input text based on a pre-built LLM) to generate a summary, and generating an emotional element analysis table in which the summary corresponding to each of the plurality of paragraphs and the emotional elements corresponding to each paragraph correspond to each other and are recorded, and then transmitting the generated emotional element analysis table to the electronic device.
[0130] For the above, a method of operation of a service providing server according to an embodiment of the present invention has been described with reference to FIG. 2. Here, since the method of operation of a service providing server according to an embodiment of the present invention may correspond to the configuration of the operation of a service providing server (110) described using FIG. 1, a more detailed description thereof will be omitted.
[0131] A method of operation of a service providing server according to one embodiment of the present invention may be implemented as a computer program stored in a storage medium for execution through combination with a computer.
[0132] In addition, the method of operation of a service providing server according to one embodiment of the present invention may be implemented in the form of computer program instructions for execution through combination with a computer and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either individually or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the present invention, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.
[0133] As described above, the present invention has been explained by specific details such as specific components, limited embodiments, and drawings; however, this is provided merely to aid in a more comprehensive understanding of the invention, and the invention is not limited to the above embodiments. A person skilled in the art can make various modifications and variations from this description.
[0134] Therefore, the scope of the present invention is not limited to the described embodiments, and all things equivalent to or having equivalent variations to the claims set forth below, as well as the claims set forth below, shall be considered to fall within the scope of the concept of the present invention. Explanation of the symbols
[0135] 110: Service Provider Server 111: Model Database 112: Matching section 113: Generation section 114: Transmission section 115: Word Database 116: Word set database 117: Word set generation section 118: Emotion Element Matching Section 119: Summary Generation Section 120: Table transfer section 10: Electronic devices
Claims
Claim 1 A service providing server comprising: a model database storing data for a speech synthesis model corresponding to each of a plurality of emotional elements—wherein the speech synthesis model corresponding to each of the plurality of emotional elements is a pre-constructed speech synthesis model based on a voice expressing each emotional element; a matching unit that, upon receiving an electronic document containing a plurality of paragraphs from a user’s electronic device and receiving a command instructing to perform speech synthesis for the electronic document, performs word analysis for each of the plurality of paragraphs inserted in the electronic document and matches one of the plurality of emotional elements corresponding to each paragraph to each of the plurality of paragraphs; a generating unit that, for each of the plurality of paragraphs, loads data for a speech synthesis model corresponding to the emotional element matched to each paragraph from the model database and performs a process of processing speech synthesis based on the loaded speech synthesis model to generate a synthesized voice corresponding to each of the plurality of paragraphs; and a transmission unit that generates a voice file in which the synthesized voices for each of the plurality of paragraphs are sequentially connected and then transmits the generated voice file to the electronic device. Claim 2 In claim 1, the matching unit comprises: a word database storing a plurality of words and a word vector corresponding to each of the plurality of words—wherein the word vector corresponding to each of the plurality of words is a pre-embedded vector based on semantic similarity between words, wherein a higher vector similarity is calculated as the semantic similarity between words increases—; a word set database storing a word set corresponding to each of the plurality of emotion elements—wherein the word set corresponding to each of the plurality of emotion elements is a set composed of two or more words pre-designated as corresponding to each emotion element among the plurality of words; and a word set generation unit that generates a word set corresponding to each of the plurality of paragraphs by performing a process of extracting words that match the plurality of words existing in each paragraph and configuring them into a word set for each of the plurality of paragraphs. A service providing server comprising an emotion element matching unit that matches an emotion element corresponding to each of the plurality of paragraphs by performing a process of, for each of the plurality of paragraphs, calculating a similarity based on word vectors between a word set corresponding to each of the plurality of emotion elements and a word set corresponding to each paragraph, and then matching one emotion element with the maximum similarity among the plurality of emotion elements to the emotion element corresponding to the paragraph. Claim 3 In paragraph 2, the emotion element matching unit performs a process of sequentially matching emotion elements for each of the plurality of paragraphs, wherein when it is the turn to perform the process of matching emotion elements for the first paragraph, which is one of the plurality of paragraphs, the unit refers to the word database to confirm the word vectors of the words constituting the word set corresponding to each of the plurality of emotion elements and the word vectors of the words constituting the word set corresponding to the first paragraph, and then for each of the plurality of emotion elements, all vector similarities between each word constituting the word set corresponding to each emotion element and each word constituting the word set corresponding to the first paragraph are calculated, and when the calculated vector similarities are sorted in ascending order and quartiles are calculated, the unit performs a process of calculating the overall average of the remaining vector similarities, excluding the vector similarities corresponding to the first quartile or lower, as the similarity corresponding to each emotion element, thereby calculating the similarity between the word set corresponding to each of the plurality of emotion elements and the word set corresponding to the first paragraph, and among the plurality of emotion elements, the emotion element with the maximum similarity is the A service providing server characterized by matching to an emotional element corresponding to the first paragraph. Claim 4 A service providing server according to claim 1, wherein the transmission unit, when generating a voice file in which synthesized voices for each of the plurality of paragraphs are sequentially connected, includes a silent interval of a preset time between the synthesized voices corresponding to each paragraph during the process of sequentially connecting the synthesized voices corresponding to each of the plurality of paragraphs, thereby generating a voice file in which a silent interval is included between the synthesized voices corresponding to each paragraph. Claim 5 A service providing server further comprising: a summary generation unit that, after a voice file is transmitted to the electronic device through the transmission unit, when a request command to provide an emotion element analysis table is received from the electronic device that analyzes and records emotion elements for each of the plurality of paragraphs, inputs the text contained in each of the plurality of paragraphs as input to a pre-set summary generation model—the summary generation model is a generative artificial intelligence model that receives text input and automatically generates a summary corresponding to the input text based on a pre-built LLM (Large Language Model)—to generate a summary; and a table transmission unit that generates an emotion element analysis table in which the summary corresponding to each of the plurality of paragraphs and the emotion elements corresponding to each paragraph correspond to each other and records them, and then transmits the generated emotion element analysis table to the electronic device. Claim 6 A method of operation of a service providing server comprising: a step of maintaining a model database in which data regarding a speech synthesis model corresponding to each of a plurality of emotional elements is stored—the speech synthesis model corresponding to each of the plurality of emotional elements is a pre-constructed speech synthesis model based on a voice expressing each emotional element—when an electronic document containing a plurality of paragraphs is received from a user’s electronic device and a command instructing to perform speech synthesis for the electronic document is received, a step of performing word analysis on each of the plurality of paragraphs inserted in the electronic document to match one of the plurality of emotional elements corresponding to each paragraph to each of the plurality of paragraphs; a step of generating a synthesized voice corresponding to each of the plurality of paragraphs by loading data regarding a speech synthesis model corresponding to the emotional element matched to each paragraph from the model database for each of the plurality of paragraphs and performing a process of processing speech synthesis based on the loaded speech synthesis model; and a step of generating a voice file in which the synthesized voices for each of the plurality of paragraphs are sequentially connected and then transmitting the generated voice file to the electronic device. Claim 7 In claim 6, the step of matching any one of the aforementioned emotional elements comprises: maintaining a word database in which a plurality of words and a word vector corresponding to each of the plurality of words are stored—the word vector corresponding to each of the plurality of words is a pre-embedded vector based on semantic similarity between words, specified such that a higher vector similarity is produced as the semantic similarity between words increases—; maintaining a word set database in which a word set corresponding to each of the plurality of emotional elements is stored—the word set corresponding to each of the plurality of emotional elements is a set composed of two or more words pre-designated as corresponding to each emotional element among the plurality of words; and for each of the plurality of paragraphs, generating a word set corresponding to each of the plurality of paragraphs by performing a process of extracting words that match the plurality of words existing in each paragraph and composing them into a word set. A method of operation of a service providing server comprising the step of matching emotional elements corresponding to each of the plurality of paragraphs by performing a process of calculating, for each of the plurality of paragraphs, a similarity based on word vectors between a word set corresponding to each of the plurality of emotional elements and a word set corresponding to each paragraph, and then matching one emotional element with the maximum similarity among the plurality of emotional elements as the emotional element corresponding to the paragraph. Claim 8 In claim 7, the step of matching the corresponding emotional elements is to sequentially perform the process of matching emotional elements for each of the plurality of paragraphs, wherein when it is the turn to perform the process of matching emotional elements for the first paragraph, which is any one of the plurality of paragraphs, the word vectors of the words constituting the word set corresponding to each of the plurality of emotional elements and the word vectors of the words constituting the word set corresponding to the first paragraph are confirmed by referring to the word database, and then for each of the plurality of emotional elements, all vector similarities between each word constituting the word set corresponding to each emotional element and each word constituting the word set corresponding to the first paragraph are calculated, and when the calculated vector similarities are sorted in ascending order and quartiles are calculated, the overall average of the remaining vector similarities, excluding the vector similarities corresponding to the first quartile or lower, is calculated as the similarity corresponding to each emotional element, thereby calculating the similarity between the word set corresponding to each of the plurality of emotional elements and the word set corresponding to the first paragraph, and among the plurality of emotional elements, any one emotional element with the maximum similarity A method of operation of a service providing server characterized by matching an element to an emotional element corresponding to the first paragraph above. Claim 9 A method of operation of a service providing server, wherein, in the transmission step, when generating a voice file in which synthesized voices for each of the plurality of paragraphs are sequentially connected, a silent interval of a preset time is included between the synthesized voices corresponding to each paragraph during the process of sequentially connecting the synthesized voices corresponding to each of the plurality of paragraphs, thereby generating a voice file in which a silent interval is included between the synthesized voices corresponding to each paragraph. Claim 10 A method of operation of a service providing server according to claim 6, further comprising: a step of generating a summary by inputting text contained in each of the plurality of paragraphs as input to a generative artificial intelligence model that receives text input and automatically generates a summary corresponding to the input text based on a pre-built LLM (Large Language Model) after a voice file is transmitted to the electronic device through the transmission step, and when a request command for providing an emotional element analysis table is received from the electronic device, which analyzes and records emotional elements for each of the plurality of paragraphs; and a step of generating an emotional element analysis table in which the summary corresponding to each of the plurality of paragraphs and the emotional elements corresponding to each paragraph correspond to each other, and then transmitting the generated emotional element analysis table to the electronic device. Claim 11 A computer-readable recording medium having a computer program for executing the method of any one of paragraphs 6 through 10 in combination with a computer. Claim 12 A computer program stored on a storage medium for executing the method of any one of paragraphs 6 through 10 through combination with a computer.