Intelligent broadcast method, system and equipment for listening books and medium

By using big model technology to identify and express the emotional characteristics of books in the e-book broadcasting system, the problem that traditional broadcasting methods are difficult to accurately express the emotions of books is solved, and a personalized and interactive listening experience is achieved.

CN120220643APending Publication Date: 2025-06-27UNICOM WOYUEDU TECH CULTURE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510214437.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional e-book broadcasting methods are difficult to accurately express the emotional characteristics of books, resulting in monotonous listening experience and inability to accurately understand the author's emotions.

Method used

Using big model technology, multiple texts are extracted from the electronic books to be broadcast, combined to generate the text to be broadcasted, and the emotional characteristics in the text are identified through the big model, and the tone, speech speed and tone of the broadcast are adjusted to accurately express emotions.

Benefits of technology

It realizes the accurate identification and expression of the emotional characteristics of e-books, provides a personalized and interactive listening experience, and improves users' appreciation and participation in literary works.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220643A_ABST
    Figure CN120220643A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent broadcasting method, system and device for listening books and a medium. The method comprises the following steps: acquiring an electronic book to be broadcasted; extracting a plurality of texts from the electronic book to be broadcasted; generating a to-be-broadcasted text based on the plurality of text combinations; the large model recognizes a first emotional feature from the text to be broadcasted; broadcasting the to-be-broadcasted text based on the first emotion feature; by acquiring the electronic book to be broadcasted, extracting a plurality of texts, combining to generate the text to be broadcasted, identifying the emotional characteristics based on a large model, and broadcasting according to the emotional characteristics, the problem that the emotional characteristics of the book are difficult to accurately express by a traditional broadcasting method is solved, and the method has the advantages of accurately expressing the emotional characteristics of the book and improving the broadcasting efficiency. The method has the advantages of providing personalized and interactive book listening experience, effectively extracting and integrating image-text contents, and effectively organizing and combing long-length or complex-structure electronic book contents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of book broadcasting, and in particular, to an intelligent broadcasting method, system, device and medium for listening to books. Background Art

[0002] An e-book, also known as an electronic book or digital book, is a new form of book that emerged with the development of electronic publishing, the Internet, and modern communication and electronic technologies. It is different from traditional publications on paper and represents digital publications that people read. The content of e-books is mainly made in a special format, can be transmitted over wired or wireless networks, and is generally organized by specialized websites.

[0003] During the process of broadcasting e-books, usually the e-books are converted into text format and then the text is played, but the prior art is difficult to accurately express the emotions of the books. Summary of the Invention

[0004] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.

[0005] The main purpose of the embodiments of the present disclosure is to propose an intelligent broadcasting method, system, device and medium for listening to books, which can solve the problem that traditional broadcasting methods are difficult to accurately express the emotional characteristics of books.

[0006] The first aspect of the embodiments of the present application proposes an intelligent broadcasting method for listening to books, and the method includes:

[0007] Obtain the e-book to be broadcast;

[0008] Extract multiple texts from the e-book to be broadcast;

[0009] Generate the text to be broadcast based on the combination of multiple texts;

[0010] Identify the first emotional feature from the text to be broadcast based on a large model;

[0011] Broadcast the text to be broadcast based on the first emotional feature.

[0012] The intelligent broadcasting method for listening to books provided by the embodiments of the present disclosure has at least the following beneficial effects:

[0013] By obtaining the e-book to be broadcasted, extracting multiple texts, combining them to generate the text to be broadcasted, identifying the emotional features based on a large model, and performing the broadcast according to the emotional features, it solves the problem that traditional broadcast methods are difficult to accurately express the emotional features of books. It has the advantages of being able to accurately express the emotional features of books, providing a personalized and interactive audiobook experience, effectively extracting and integrating graphic and text content, and effectively organizing and sorting out the content of long or structurally complex e-books.

[0014] In some embodiments, the extracting multiple texts from the e-book to be broadcasted includes:

[0015] Determining multiple text regions from the e-book to be broadcasted;

[0016] Based on image-to-text technology, identifying multiple texts in the multiple text regions.

[0017] In some embodiments, the combining multiple texts to generate the text to be broadcasted includes:

[0018] Obtaining the keyword of the e-book; the keyword is a word or sentence that can characterize the basic information of the e-book;

[0019] According to the keyword, organizing the multiple texts into the text to be broadcasted.

[0020] In some embodiments, the large model is Chat-Gpt or DeepSeek.

[0021] In some embodiments, before identifying the first emotional feature from the text to be broadcasted based on the large model, the method further includes:

[0022] Identifying the physical sign information of the user;

[0023] Generating a second emotional feature based on the physical sign information of the user;

[0024] The performing the broadcast of the text to be broadcasted based on the first emotional feature includes:

[0025] Performing the broadcast of the text to be broadcasted according to the first emotional feature and the second emotional feature.

[0026] In some embodiments, the physical sign information includes any one of facial features and heart rate features.

[0027] In some embodiments, both the first emotional feature and the second emotional feature include any one of neutral, joy, sadness, anger, sleepiness, disgust, surprise or fear.

[0028] In a second aspect of the embodiments of the present application, an intelligent playback system for audiobooks is proposed. The system includes:

[0029] A book acquisition unit for acquiring the e-books to be played;

[0030] A text extraction unit for extracting multiple texts from the e-books to be played;

[0031] A text acquisition unit for generating the text to be played based on the combination of multiple texts;

[0032] An emotion feature acquisition unit for identifying the first emotion feature from the text to be played based on a large model;

[0033] A text playback unit for playing the text to be played based on the first emotion feature.

[0034] In a third aspect of the embodiments of the present application, an electronic device is proposed, including at least one controller and a memory communicatively connected to the at least one controller; the memory stores instructions executable by the at least one controller, and when the instructions are executed by the controller, the controller is caused to execute the intelligent playback method for audiobooks as described in the first aspect.

[0035] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is proposed. The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the intelligent playback method for audiobooks as described above.

[0036] It can be understood that the beneficial effects of the above second aspect to the fourth aspect compared with the related art are the same as those of the above first aspect compared with the related art. For the relevant descriptions, reference can be made to the relevant descriptions in the first aspect, and details will not be repeated here. Description of the Drawings

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the related art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 is a schematic diagram of an intelligent playback method for audiobooks provided by the present application

[0039] Figure 2 is a schematic diagram of an intelligent playback system for audiobooks proposed in the embodiments of the present application;

[0040] Figure 3It is a schematic diagram of an electronic device proposed in an embodiment of the present application. Detailed implementation manners

[0041] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0042] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0044] As a new form of digital publication, e-books have become an important carrier for modern reading and information acquisition. However, during the broadcast of e-books, there are technical problems in accurately expressing the emotions of the books. This problem mainly stems from the current method of simply converting e-books into text format for playback, which cannot effectively capture and convey the emotional characteristics contained in the original text.

[0045] Specifically, in an e-book broadcast system, the text-to-speech conversion process usually adopts standardized speech synthesis technology. Although this technology can accurately convert text into speech, it often lacks the ability to recognize and express the emotional characteristics of the text content. For example, when broadcasting a text describing a sad scene, the system may broadcast it in a flat or inappropriate tone, resulting in the emotions conveyed by the original text not being accurately presented to the listeners. In addition, since the content of e-books often contains complex plots and rich emotional changes, a single speech synthesis method is difficult to adapt to this diversity, thus affecting the user's audiobook experience.

[0046] Therefore, if this technical problem cannot be effectively solved, it will have a significant negative impact on the usability and user experience of the e-book broadcasting system. First of all, listeners may not be able to accurately understand the emotions that the author wants to convey, resulting in misunderstandings or reduced appreciation of literary works. Secondly, the lack of emotional expression in the broadcasting method may make the process of listening to books monotonous and boring, reducing user participation and willingness to continue using. In addition, for certain types of e-books, such as children's books or emotional novels, accurate emotional expression is particularly important. Failing to solve this problem will seriously limit the broadcasting effect of these types of e-books. Therefore, developing an intelligent broadcasting method that can accurately identify and express the emotional characteristics of e-books is of great significance for improving the technical level and user experience of the e-book broadcasting system.

[0047] In view of the technical problem that it is difficult to accurately express the emotions of books during the e-book broadcasting process, this application has conducted in-depth thinking and exploration.

[0048] First of all, considering the complexity of e-book content and the diversity of emotional expression, a method that can comprehensively understand the text content is needed. Among them, the large model technology has become a possible solution due to its powerful natural language processing ability. The large model can understand the semantics and context of the text, and thus may be able to identify the emotional characteristics contained in the text.

[0049] Specifically, this application proposes a method of using a large model to identify emotional characteristics from the text to be broadcast. The advantage of this method is that the large model can perform in-depth semantic analysis on the text based on a large amount of training data, so as to accurately capture the emotional information in the text. For example, when encountering text describing a sad scene, the large model can identify the sad emotion in it, providing emotional guidance for subsequent broadcasting.

[0050] However, simply identifying the emotional characteristics is not enough to solve the problem. How to apply the identified emotional characteristics to the actual broadcasting process becomes the next problem to be solved. Therefore, this application further proposes a method of broadcasting the text to be broadcast based on the identified emotional characteristics. This method can adjust the tone, speed, and pitch of the broadcast according to different emotional characteristics, so as to more accurately express the emotions in the text.

[0051] Such as Figure 1 , in order to implement this method, this application designs a complete processing flow.

[0052] Step S110, obtain the e-book to be broadcast;

[0053] Step S120, extract multiple texts from the e-book to be broadcast;

[0054] Step S130, generate the text to be broadcast based on the combination of multiple texts;

[0055] Step S140, identifying a first emotional feature from the text to be broadcast based on a large model;

[0056] Step S150, broadcasting the text to be broadcast based on the first emotional feature.

[0057] This method can more accurately extract the text content from e-books by first determining the text area and then using image-to-text technology for recognition. This method is particularly suitable for e-books with mixed text and images, and can effectively distinguish the text content from non-text content such as pictures.

[0058] In practical applications, various methods can be used to determine the text area. For example, image segmentation algorithms can be used to divide the e-book page into different areas, and then the text area can be identified through feature analysis. Or, a machine learning model can be used to train a classifier that can automatically identify the text area.

[0059] During the process of identifying text based on image-to-text technology, optical character recognition (OCR) technology can be used. OCR technology can convert the text in the image into an editable text format. Specifically, when implementing, an OCR algorithm suitable for the characteristics of e-books can be selected, such as specialized algorithms for different fonts and different languages, to improve the accuracy of recognition.

[0060] In some cases, e-books may contain complex layout or special fonts. To address this situation, the method of this application can further combine layout analysis technology. Layout analysis can help understand the layout structure of the text area, so as to more accurately extract the text content. For example, different types of text areas such as titles, main texts, headers and footers can be identified, and different processing strategies can be adopted according to their characteristics.

[0061] In addition, to improve the efficiency and accuracy of text extraction, the method of this application can also adopt parallel processing technology. For example, different pages or different areas of the e-book can be processed simultaneously, making full use of computing resources to speed up the text extraction speed.

[0062] As a specific embodiment, assume that there is an e-book containing 300 pages that needs text extraction. First, use an image segmentation algorithm to divide each page into several regions. Then, apply a pre-trained deep learning model to identify the text regions in these regions. After identifying the text regions, use an OCR algorithm optimized for Chinese to recognize the text content. During this process, the system will process multiple pages simultaneously, for example, 10 pages at a time. For the recognized text, the system will also perform post-processing, such as removing extra spaces, correcting common recognition errors, etc. Finally, the system combines all the recognized texts in the original order of the book to form the complete text content.

[0063] Compared with the prior art, the method proposed in this application has the following advantages: First, by using a two-step method of first determining the text region and then performing text recognition, the accuracy of text extraction is improved, especially for e-books with complex layouts; Second, by adopting image-to-text technology, the method is applicable to e-books in various formats and is not limited to a specific file format; Finally, through technical means such as parallel processing, the efficiency of text extraction is improved. These advantages enable the method of this application to better solve the problem of e-book text extraction and provide a high-quality text basis for subsequent intelligent broadcasting.

[0064] In some of the above embodiments, during the implementation of this application, there is also a problem of how to more effectively combine multiple texts to generate the text to be broadcast.

[0065] In response to this, this application further proposes a method for generating the text to be broadcast based on the combination of multiple texts, including obtaining the keywords of the e-book. The keywords are words or sentences that can represent the basic information of the e-book, and then organizing multiple texts into the text to be broadcast according to the keywords.

[0066] This method provides an effective guiding direction for text combination by introducing the concept of keywords. As representatives of the basic information of e-books, keywords can help the system better understand and organize text content.

[0067] Specifically, first obtain the keywords of the e-book. These keywords can be words or short sentences representing the theme, characters, plot, or main concepts of the e-book. For example, for a science fiction novel, the keywords may include "space exploration", "artificial intelligence", "time travel", etc.

[0068] There are various methods for obtaining keywords. One way is to extract them from the table of contents, abstract, or text of the e-book through natural language processing techniques. Another way is to utilize metadata information, such as the tags or classification information of the book. It is also possible to combine artificial intelligence techniques and use machine learning models to identify and extract keywords.

[0069] Next, the system will organize multiple texts based on these keywords. This process may include the following steps:

[0070] 1. Text relevance analysis: The system will evaluate the degree of relevance of each text segment to the keywords. This can be achieved by calculating factors such as the frequency of occurrence and position of the keywords in the text.

[0071] 2. Text sorting: Based on the results of the relevance analysis, the system will sort the texts, placing the texts most relevant to the keywords at the front.

[0072] 3. Text combination: The system will combine the text segments together according to the sorting results to form a coherent text to be broadcast. In this process, some transitional statements may need to be added to ensure the fluency of the text.

[0073] 4. Content optimization: The system may perform further optimizations, such as removing duplicate content, adjusting the paragraph order, etc., to improve the quality of the text to be broadcast.

[0074] Through this method, this application can more effectively organize and arrange the text content in e-books, generating a text to be broadcast with a reasonable structure and prominent key points. This not only improves the coherence and readability of the text but also ensures that the broadcast content better reflects the core information of the e-book.

[0075] As a preferred implementation, the following specific examples can be considered:

[0076] Suppose there is an e-book named "Artificial Intelligence and Future Society". The system first extracts keywords such as "artificial intelligence", "machine learning", "social impact", "ethical issues", etc. by analyzing the table of contents and abstract of the book.

[0077] Then, the system extracts multiple text segments from the e-book. For example:

[0078] Segment 1: "Artificial intelligence technology is developing rapidly and affecting all aspects of our lives."

[0079] Segment 2: "Machine learning algorithms can learn patterns and rules from large amounts of data."

[0080] Segment 3: "With the popularization of AI, we need to consider its impact on the job market."

[0081] Segment 4: "Artificial intelligence shows great potential in the field of medical diagnosis."

[0082] The system will perform relevance analysis and sorting on these segments based on keywords. In this example, Segment 1 and Segment 2 are highly relevant to the keywords "artificial intelligence" and "machine learning" and may be ranked at the front. Segment 3, which involves "social impact", will also be included in the text to be broadcast. Although Segment 4 is related to artificial intelligence, its relevance to the main keywords may be relatively low, and it may be ranked behind or omitted.

[0083] Finally, the system will combine these segments into a coherent text to be broadcast, which may be as follows:

[0084] "Artificial intelligence technology is developing rapidly and affecting all aspects of our lives. Among them, machine learning algorithms can learn patterns and rules from a large amount of data, which is one of the key technologies in the development of artificial intelligence. With the popularization of AI, we need to consider its impact on the job market, which is a social issue that cannot be ignored in the process of artificial intelligence development."

[0085] Compared with simply combining texts in the order they appear in the book, this method can better highlight the core content of the e-book, making the generated text to be broadcast more refined and focused. At the same time, through the guidance of keywords, it can ensure that the content between different chapters or paragraphs is more coherent, improving the quality of the audiobook experience.

[0086] In addition, this method also has strong flexibility and scalability. For example, the weights of keywords can be adjusted according to the user's interests or the purpose of listening to the book, so as to generate a text to be broadcast that better meets the user's needs. It can also combine natural language processing technology to further optimize the text combination logic, such as considering factors like semantic coherence and emotional consistency.

[0087] Compared with the prior art, the method of the present application has obvious advantages in text combination. Traditional methods may simply combine texts in the order they appear in the book or use some basic rules to select texts. However, the present application provides a more intelligent and targeted method for text combination by introducing the concept of keywords. This not only improves the quality of the text to be broadcast but also better retains and conveys the core information of the e-book, thus significantly improving the user's audiobook experience.

[0088] In some of the above embodiments, during the implementation of the present application, there is also a problem of being unable to accurately identify the emotional characteristics in the text to be broadcast.

[0089] In response to this, the present application further proposes a technical solution where the large model is Chat-Gpt or DeepSeek.

[0090] The intelligent broadcast method for audiobooks proposed in this application includes the following steps: obtaining an e-book to be broadcast; extracting multiple texts from the e-book to be broadcast; generating a text to be broadcast based on the combination of multiple texts; identifying a first emotional feature from the text to be broadcast based on a large model; and broadcasting the text to be broadcast based on the first emotional feature. Among them, extracting multiple texts from the e-book to be broadcast includes: determining multiple text areas from the e-book to be broadcast; and identifying multiple texts in the multiple text areas based on image-to-text technology. Generating a text to be broadcast based on the combination of multiple texts includes: obtaining a keyword of the e-book, where the keyword is a word or sentence that can represent the basic information of the e-book; and sorting the multiple texts into a text to be broadcast according to the keyword.

[0091] In this application, by using Chat-Gpt or DeepSeek as the large model to identify the emotional features in the text to be broadcast, the emotional information in the text can be captured more accurately. Both of these large models have powerful natural language processing capabilities, can understand the context and semantics of the text, and thus can more precisely identify the emotions contained in the text.

[0092] Specifically, Chat-Gpt is a large-scale language model based on the transformer architecture. It has learned rich language knowledge and context understanding capabilities through pre-training on a large amount of text data. When identifying emotional features, Chat-Gpt can analyze the language structure, word selection, and tone of the text, and thus infer the emotions expressed by the text.

[0093] DeepSeek is another advanced natural language processing model, which also has powerful text understanding and analysis capabilities. DeepSeek can extract emotion-related features from the text through deep learning algorithms and map these features to different emotion categories.

[0094] Using either of these two large models can significantly improve the accuracy of emotional feature recognition. For example, when encountering a text describing a sad scene, the large model can accurately identify the sad emotion in the text by analyzing the keywords, sentence structure, and overall context in the text. This accurate emotion recognition provides an important basis for subsequent broadcasts, enabling the broadcast to better convey the emotional content of the text.

[0095] Furthermore, since both Chat-Gpt and DeepSeek are models that are continuously updated and optimized, their performance will continuously improve over time. This means that the method of this application can automatically obtain performance improvement as the large model progresses, without the need to frequently update the system.

[0096] As a preferred embodiment, Chat-Gpt can be selected as the large model. In specific implementation, first, the text to be broadcast is input into the Chat-Gpt model. Then, by setting appropriate prompts, Chat-Gpt is required to analyze the emotional characteristics of the text. For example, the following prompt can be used: "Please analyze the emotional characteristics of the following text and give the main emotional types." Next, Chat-Gpt will output the analysis results, including the identified main emotional types (such as joy, sadness, anger, etc.) and the corresponding confidence levels. Finally, the system selects appropriate speech synthesis parameters (such as intonation, speech rate, volume, etc.) for broadcasting based on this analysis result, so as to better express the emotional content of the text.

[0097] Compared with the prior art, the application of using Chat-Gpt or DeepSeek as the large model to identify emotional characteristics has significant advantages. Traditional emotion recognition methods usually rely on predefined emotion lexicons or simple rules and are difficult to accurately capture complex emotional expressions. However, the present application utilizes the powerful language understanding ability of the large model to more comprehensively analyze the semantics and context of the text, thereby achieving more accurate emotion recognition. This method can not only identify basic emotion types but also capture more delicate emotional changes, bringing a qualitative improvement to the audiobook experience.

[0098] In some of the above embodiments, during the implementation of the present application, there is still a problem of difficultly accurately identifying the user's emotional state.

[0099] In response to this, the present application further proposes a technical solution in which the physical sign information includes any one of facial features and heart rate features.

[0100] The technical solution of the present application more accurately identifies the user's emotional state by introducing the user's physical sign information, including facial features and heart rate features. This method combines physiological and psychological knowledge and can provide a more comprehensive and objective emotion assessment.

[0101] Specifically, the technical solution of the present application can be implemented in the following ways:

[0102] 1. Facial feature recognition: Use computer vision technology, such as deep learning algorithms, to analyze the user's facial expressions. This may include detecting minute changes such as the position of the eyebrows, the curvature of the mouth corners, and the degree of eye opening and closing. For example, a smile may indicate joy, and a frown may indicate confusion or dissatisfaction.

[0103] 2. Heart rate feature analysis: Monitor the user's heart rate changes through wearable devices or non-contact sensors. The speed and change pattern of the heart rate can reflect the user's emotional state. For example, an accelerated heart rate may indicate excitement or tension, while a stable heart rate may indicate calmness or concentration.

[0104] 3. Data fusion: Integrate the data of facial features and heart rate features for comprehensive analysis, and use machine learning algorithms to establish a more accurate emotion recognition model. This multi-modal approach can improve the accuracy and robustness of emotion recognition.

[0105] 4. Real-time feedback: Based on the recognized emotion state, the system can dynamically adjust the tone, speed, and pitch of the broadcast content to better match the user's emotional needs.

[0106] Through this method, the technical solution of this application can more accurately capture the user's emotional changes, thereby providing a more personalized and adaptable audiobook experience. For example, when it is detected that the user has a fast heart rate and a tense facial expression, the system may choose a more gentle intonation to broadcast the content to help the user relax.

[0107] As a preferred implementation, an infrared camera can be used to capture the user's facial heat map, combined with facial expression analysis algorithms, to more accurately identify facial features. At the same time, a sensor using photoplethysmography (PPG) technology can be used to non-invasively monitor the user's heart rate changes. The combination of these technologies can continuously collect data on the user's emotional state without disturbing the user's normal audiobook experience.

[0108] The specific embodiments are as follows:

[0109] In an intelligent audiobook device, a high-definition camera and an infrared sensor are integrated. When the user starts using the audiobook function, the device first performs user authorization confirmation. After obtaining authorization, the device starts to collect the user's facial images and heart rate data. The facial images are analyzed through a deep learning model to extract key facial feature points, such as the positions and shape changes of eyebrows, eyes, and the corners of the mouth. At the same time, the heart rate data is processed through signal processing techniques such as Fourier transform to extract heart rate variability (HRV) indicators.

[0110] These data are input into a pre-trained multi-modal emotion recognition model. The model adopts an attention mechanism and a long short-term memory network (LSTM), which can comprehensively consider the temporal changes of facial features and heart rate features. The model outputs the user's current emotional state, such as neutral, joy, sadness, etc.

[0111] Based on the recognized emotion state, the audiobook system will adjust the broadcast parameters in real time. For example, when it is recognized that the user's emotion is sadness, the system may choose a more gentle voice synthesis tone and slightly reduce the broadcast speed to give the user more time to understand and feel the content.

[0112] Compared with the prior art, the technical solution of this application has the following advantages:

[0113] 1. Multimodal Fusion: By combining facial features and heart rate features, it provides more comprehensive and accurate emotion recognition than single modality.

[0114] 2. Non-invasive: Adopting non-contact sensing technology, it avoids the discomfort that traditional emotion recognition methods may bring.

[0115] 3. Real-time: It can continuously monitor and analyze the user's emotional state, achieve dynamic adjustment, and provide a more personalized audiobook experience.

[0116] 4. Strong adaptability: This method can adapt to the individual differences of different users and improve the recognition accuracy through continuous learning.

[0117] In this way, the technical solution of this application effectively solves the problem in the prior art that it is difficult to accurately capture the user's emotional changes, and provides a more user-friendly and intelligent solution for the intelligent audiobook system.

[0118] Such as Figure 2 , an embodiment of this application provides an intelligent playback system for audiobooks, and the system includes:

[0119] The book acquisition unit 1100 is used to acquire the e-books to be played;

[0120] The text extraction unit 1200 is used to extract multiple texts from the e-books to be played;

[0121] The text acquisition unit 1300 is used to generate the text to be played based on the combination of multiple texts;

[0122] The emotion feature acquisition unit 1400 is used to identify the first emotion feature from the text to be played based on the large model;

[0123] The text playback unit 1500 is used to play the text to be played based on the first emotion feature.

[0124] It should be noted that an intelligent playback system for audiobooks provided in this embodiment and the above-mentioned intelligent playback method for audiobooks are based on the same inventive concept. Therefore, the relevant content of the above-mentioned intelligent playback method for audiobooks also applies to the content of the intelligent playback system for audiobooks. Therefore, it will not be elaborated here.

[0125] Such as Figure 3 , an embodiment of this application also provides an electronic device, and this electronic device includes:

[0126] At least one memory;

[0127] At least one processor;

[0128] At least one program;

[0129] The program is stored in the memory, and the processor executes at least one program to implement an intelligent broadcast method for listening to books as described above in the present disclosure.

[0130] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.

[0131] The following details the electronic device according to the embodiments of the present application.

[0132] The processor 1600 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention.

[0133] The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700 and are called by the processor 1600 to execute an intelligent broadcast method for listening to books according to the embodiments of the present invention.

[0134] The input / output interface 1800 is used to implement information input and output.

[0135] The communication interface 1900 is used to implement communication interaction between this device and other devices. It can communicate through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as a mobile network, WIFI, Bluetooth, etc.).

[0136] The bus 2000 transmits information between various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900).

[0137] Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.

[0138] An embodiment of the present invention further provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-mentioned intelligent playback method for audiobooks.

[0139] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices.

[0140] In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0141] The embodiments described in the present invention are for more clearly illustrating the technical solutions of the embodiments of the present invention and do not constitute a limitation to the technical solutions provided by the embodiments of the present invention. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.

[0142] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present invention, and may include more or fewer steps than those shown, or combine some steps, or different steps.

[0143] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0144] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0145] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0146] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0147] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0148] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0149] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0150] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store programs.

[0151] The above is a specific description of the preferred implementation of the embodiments of the present application. However, the embodiments of the present application are not limited to the above implementation manners. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the embodiments of the present application. These equivalent deformations or substitutions are all included within the scope defined by the claims of the embodiments of the present application.

Claims

1. An intelligent broadcasting method for listening to books, characterized in that: The method comprises: Get the electronic books to be broadcast; Extracting a plurality of texts from the electronic book to be broadcasted; Generate a text to be broadcast based on a combination of multiple texts; Identifying a first emotion feature from the text to be broadcast based on the large model; The text to be announced is announced based on the first emotion feature.

2. The intelligent broadcasting method for listening to books according to claim 1, characterized in that: The step of extracting a plurality of texts from the electronic book to be broadcast comprises: Determining a plurality of text areas from the electronic book to be broadcast; Based on the image-to-text technology, multiple texts in multiple text areas are identified.

3. The intelligent broadcasting method for listening to books according to claim 2 is characterized in that: The generating of the to-be-broadcasted text based on the combination of multiple texts includes: Acquire keywords of the electronic book; the keywords are words or sentences that can represent basic information of the electronic book; According to the keyword, the multiple texts are sorted into the text to be broadcast.

4. The intelligent broadcasting method for listening to books according to claim 3 is characterized in that: The large model is Chat-Gpt or DeepSeek.

5. The intelligent broadcasting method for listening to books according to claim 4 is characterized in that: Before identifying the first emotion feature from the to-be-broadcasted text based on the large model, the method further includes: Identify the user's vital signs; generating a second emotion feature based on the user's physical sign information; The reporting the text to be reported based on the first emotion feature includes: The text to be broadcast is broadcasted according to the first emotion feature and the second emotion feature.

6. The intelligent broadcasting method for listening to books according to claim 5, characterized in that: The physical sign information includes any one of facial features and heart rate features.

7. The intelligent broadcasting method for listening to books according to claim 6, characterized in that: The first emotion characteristic and the second emotion characteristic both include any one of neutral, joy, sadness, anger, sleepiness, disgust, surprise or fear.

8. An intelligent broadcasting system for listening to books, characterized in that: The system comprises: A book acquisition unit, used for acquiring electronic books to be broadcast; A text extraction unit, used for extracting a plurality of texts from the electronic book to be broadcast; A text acquisition unit, used for generating a text to be broadcast based on a combination of multiple texts; An emotion feature acquisition unit, configured to identify a first emotion feature from the text to be broadcast based on a large model; The text broadcasting unit is used to broadcast the text to be broadcast based on the first emotion feature.

9. An electronic device, characterized in that: include: at least one controller and a memory for communicatively coupling with the at least one controller; The memory stores instructions that can be executed by the at least one controller, and the instructions are executed by the controller so that the controller executes the intelligent broadcasting method for listening to books as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the intelligent broadcasting method for audiobooks according to any one of claims 1 to 7.