Artificial intelligence-based news broadcasting methods and related devices

CN121075311BActive Publication Date: 2026-04-03JIANGXI RONG MEDIA BRAIN TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

[0003]但是,目前大多数语音播报在处理新闻的文字和图表时,仅仅是简单地将文字和图表的内容转换为语音,容易出现图表信息遗漏或者图表信息转化生硬的问题,且缺乏对新闻内容的深入理解和情感表达,无法吸引用户的注意力,难以满足用户对新闻质量的要求

Benefits of technology

[0024]第五方面,本申请实施例提供了一种计算机程序产品,其中,上述计算机程序产品包括存储了计算机程序的非瞬时性计算机可读存储介质,上述计算机程序可操作来使计算机执行如本申请实施例第一方面任一方法中所描述的部分或全部步骤。该计算机程序产品可以为一个软件安装包。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075311B_ABST
    Figure CN121075311B_ABST
Patent Text Reader

Abstract

This application provides an artificial intelligence-based news broadcasting method and related apparatus. The method includes: acquiring text and graphic information corresponding to a target news story; processing the text and graphic information to obtain the target broadcast text and its corresponding target feature information; analyzing the target feature information to obtain reference content attributes and reference broadcast sentiment; determining the target broadcasting style based on the reference content attributes and reference broadcast sentiment; acquiring the target user's target broadcasting needs; synthesizing speech from the target broadcast text based on the target broadcasting style and target broadcasting needs using a preset AI broadcasting model to obtain the target broadcasting speech; and broadcasting the news using the target broadcasting speech in response to the target user's voice broadcasting operation. By processing the text and graphic information of the news and combining it with the user's personalized needs, and synthesizing adapted speech through an AI broadcasting model, the quality of news broadcasting can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice broadcasting technology, and in particular to a news broadcasting method and related apparatus based on artificial intelligence. Background Technology

[0002] Traditional news reading requires users to concentrate and read word by word, which can be a burden for users with limited time. Meanwhile, the development of voice technology has provided new avenues for news dissemination. By having news read aloud, users can receive information without using their hands or eyes, making it more suitable for various life scenarios, such as driving, exercising, or doing housework.

[0003] However, most current voice broadcasts simply convert text and charts into speech when processing news, which can easily lead to omissions of chart information or awkward conversion of chart information. Furthermore, they lack in-depth understanding of the news content and emotional expression, failing to attract users' attention and meet users' requirements for news quality.

[0004] Therefore, how to improve the quality of news broadcasts is an urgent issue that needs to be addressed. Summary of the Invention

[0005] This application provides an artificial intelligence-based news broadcasting method and related apparatus. By processing the text and image information of the news, analyzing its content attributes and broadcasting emotions to determine the broadcasting style, and then combining it with the user's personalized needs, the AI ​​broadcasting model synthesizes adapted voice and responds to the user's operation to complete the broadcast, thereby improving the quality of news broadcasting.

[0006] In a first aspect, embodiments of this application provide a news broadcasting method based on artificial intelligence, the method comprising:

[0007] Retrieve the text and chart information corresponding to the target news;

[0008] The text information and the chart information are processed to obtain the target broadcast text and its corresponding target feature information;

[0009] The target feature information is analyzed to obtain reference content attributes and reference broadcast sentiment.

[0010] The target broadcast style is determined based on the reference content attributes and the reference broadcast sentiment.

[0011] Obtain the target users' desired broadcasting needs;

[0012] Using a preset AI broadcasting model, the target broadcasting text is synthesized into speech based on the target broadcasting style and the target broadcasting requirements to obtain the target broadcasting speech;

[0013] In response to the target user's voice broadcast operation, the target broadcast voice is used for broadcasting.

[0014] Secondly, embodiments of this application provide an artificial intelligence-based news broadcasting device, the device comprising an acquisition module, a processing module, an analysis module, a determination module, a synthesis module, and a broadcasting module, wherein:

[0015] The acquisition module is used to acquire the text information and chart information corresponding to the target news;

[0016] The processing module is used to process the text information and the chart information to obtain the target broadcast text and its corresponding target feature information;

[0017] The analysis module is used to analyze the target feature information to obtain reference content attributes and reference broadcast sentiment.

[0018] The determining module is used to determine the target broadcasting style based on the reference content attributes and the reference broadcasting sentiment;

[0019] The acquisition module is also used to acquire the target user's target broadcasting requirements;

[0020] The synthesis module is used to synthesize the target broadcast text into speech based on the target broadcast style and the target broadcast requirements using a preset AI broadcast model, so as to obtain the target broadcast speech.

[0021] The broadcast module is used to respond to the voice broadcast operation of the target user and broadcast the target voice.

[0022] Thirdly, embodiments of this application provide an electronic device, including a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the programs include instructions for performing steps in any method of the first aspect of this application.

[0023] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in any method of the first aspect of this application.

[0024] Fifthly, embodiments of this application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in any method of the first aspect of this application. The computer program product may be a software installation package.

[0025] By implementing the embodiments of this application, the text and image information of news can be processed, its content attributes and broadcast emotions can be analyzed to determine the broadcast style, and then combined with the user's personalized needs, the AI ​​broadcast model synthesizes adapted voice and responds to the user's operation to complete the broadcast, thereby improving the quality of news broadcasting. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a system architecture diagram of a news broadcasting system provided in an embodiment of this application;

[0028] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0029] Figure 3 This is an application scenario diagram of news broadcasting provided in an embodiment of this application;

[0030] Figure 4 This is a flowchart illustrating an artificial intelligence-based news broadcasting method provided in an embodiment of this application;

[0031] Figure 5 This is a schematic diagram of a process for determining reference broadcast sentiment provided in an embodiment of this application;

[0032] Figure 6 This is a flowchart illustrating a process for determining a target broadcast style, provided in an embodiment of this application.

[0033] Figure 7 This is a schematic diagram of a feedback optimization process provided in an embodiment of this application;

[0034] Figure 8 This is a block diagram of the functional modules of a news broadcasting device based on artificial intelligence, provided in an embodiment of this application. Detailed Implementation

[0035] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0036] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0037] It should be understood that the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document indicates that the preceding and following related objects are in an "or" relationship. In the embodiments of this application, "multiple" refers to two or more.

[0038] In the embodiments of this application, "at least one item" or its similar expression refers to any combination of these items, including any combination of a single item or a plurality of items. "One or more" means one or more, while "multiple" means two or more. For example, "at least one item" of a, b, or c can represent the following seven cases: a, b, c; a and b; a and c; b and c; a, b, and c. Each of a, b, and c can be an element or a set containing one or more elements.

[0039] In this application, the term "connection" refers to various connection methods, such as direct connection or indirect connection, to achieve communication between devices. This application does not impose any limitations on this.

[0040] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0041] Traditional news reading requires users to concentrate and read word by word, which can be a burden for users with limited time. Meanwhile, the development of voice technology has provided new avenues for news dissemination. By having news read aloud, users can receive information without using their hands or eyes, making it more suitable for various life scenarios, such as driving, exercising, or doing housework.

[0042] However, most current voice broadcasts simply convert text and charts into speech when processing news, which can easily lead to omissions of chart information or awkward conversion of chart information. Furthermore, they lack in-depth understanding of the news content and emotional expression, failing to attract users' attention and meet users' requirements for news quality.

[0043] Therefore, how to improve the quality of news broadcasts is an urgent issue that needs to be addressed.

[0044] To address the aforementioned problems, this application provides an artificial intelligence-based news broadcasting method and related apparatus. First, text and graphic information corresponding to the target news are acquired. The text and graphic information are then processed to obtain the target broadcast text and its corresponding target feature information. Next, the target feature information is analyzed to obtain reference content attributes and reference broadcast sentiment. A target broadcasting style is determined based on the reference content attributes and the reference broadcast sentiment. Then, the target user's target broadcasting needs are acquired. Using a preset AI broadcasting model, the target broadcast text is synthesized into speech based on the target broadcasting style and the target broadcasting needs to obtain the target broadcasting speech. Finally, in response to the target user's voice broadcasting operation, the news is broadcast using the target broadcasting speech. By processing the text and graphic information of the news, analyzing its content attributes and broadcast sentiment to determine the broadcasting style, and then combining this with the user's personalized needs, an AI broadcasting model synthesizes adapted speech and responds to user operations to complete the broadcast, thereby improving the quality of news broadcasting.

[0045] For easier understanding, please refer to Figure 1 , Figure 1 This is a system architecture diagram of a news broadcasting system provided in an embodiment of this application. The news broadcasting system includes an information acquisition unit, a text processing unit, a feature analysis unit, a speech synthesis unit, and a feedback optimization unit.

[0046] The information collection unit is responsible for collecting raw information related to the target news. This raw information includes textual information and chart information. The textual information includes the news title, lead, body, core data description, and other textual content. The chart information includes data charts (line charts, bar charts), scene pictures (event scene pictures, people pictures), and other visual information from the target news.

[0047] The text processing unit is responsible for processing the text and chart information acquired by the information acquisition unit, converting it into the target broadcast text. Specifically, it segments the text information, extracts keywords, and performs correlation analysis to generate a logically coherent first text. Then, it identifies chart information (such as reading chart data and interpreting its meaning) and converts it into a broadcastable second text. Finally, it determines the similarity between the first and second texts, merges consistent information, generates the target broadcast text, and extracts its corresponding target feature information.

[0048] The feature analysis unit performs in-depth analysis of the target feature information output by the text processing unit, extracting key features and determining the core basis for broadcasting. This feature analysis unit can extract domain keywords, time keywords, and sentiment keywords from the target feature information. Then, it determines the reference content attributes and reference broadcast sentiment based on the keywords. Finally, combining the reference content attributes and reference broadcast sentiment, it determines the corresponding target broadcasting style.

[0049] The speech synthesis unit can acquire the target user's desired broadcasting requirements, and then synthesize the target broadcasting text into speech using a preset AI broadcasting model, combining the target broadcasting style and the target broadcasting requirements. Finally, in response to the target user's voice broadcasting operation (such as clicking the "play" button), it outputs the target broadcasting speech and adapts it to the device (such as a mobile phone, speaker, etc.) and environment (such as increasing the volume in a noisy environment).

[0050] The feedback optimization unit collects feedback data from target users on the target broadcast speech, analyzes the feedback data to obtain key information, optimization direction, and optimization intensity. Then, based on the key information, it determines the broadcast parameters that need optimization. Finally, it optimizes the broadcast parameters according to the optimization direction and intensity, and regenerates the target broadcast speech based on the optimized broadcast parameters.

[0051] It is evident that by collecting and processing news information, performing personalized voice synthesis, and continuously optimizing based on feedback data, we can ensure that the broadcast content is accurate and the style matches the news attributes, meet the personalized needs of users, and dynamically iterate to improve the experience, providing users with efficient and personalized news broadcasting services.

[0052] The following is combined Figure 2 The electronic devices in the embodiments of this application will be described. Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 2 As shown, the electronic device includes one or more processors, a memory, a communication interface, and one or more programs. The processor is connected to the memory and the communication interface via an internal communication bus.

[0053] The processor can be used for:

[0054] Retrieve the text and chart information corresponding to the target news;

[0055] The text information and the chart information are processed to obtain the target broadcast text and its corresponding target feature information;

[0056] The target feature information is analyzed to obtain reference content attributes and reference broadcast sentiment.

[0057] The target broadcast style is determined based on the reference content attributes and the reference broadcast sentiment.

[0058] Obtain the target users' desired broadcasting needs;

[0059] Using a preset AI broadcasting model, the target broadcasting text is synthesized into speech based on the target broadcasting style and the target broadcasting requirements to obtain the target broadcasting speech;

[0060] In response to the target user's voice broadcast operation, the target broadcast voice is used for broadcasting.

[0061] The one or more programs are stored in the aforementioned memory and configured to be executed by the aforementioned processor, and the one or more programs include instructions for performing any step in the above method embodiments.

[0062] The processor can be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, cells, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The communication unit can be a communication interface, transceiver, transceiver circuit, etc., and the storage unit can be a memory.

[0063] The memory can be volatile or non-volatile, or a combination of both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0064] It is understood that the electronic device may include more or fewer structural elements than those shown in the block diagram above, such as a power module, physical buttons, a Wi-Fi module, a speaker, a Bluetooth module, sensors, a display module, etc., without limitation. It is understood that the electronic device may incorporate elements such as... Figure 1 The system architecture described above.

[0065] For easier understanding, please refer to Figure 3 , Figure 3 This application provides an example of a news broadcasting scenario. The system first inputs the text information (such as title and body text) and chart information (such as data charts and scene diagrams) of the target news, along with the target user's desired broadcasting requirements, into the news broadcasting system. Then, the system processes this input information, performing operations such as text fusion, feature analysis, and speech synthesis to generate the desired broadcasting audio. Finally, the system delivers the desired broadcasting audio to the target user via voice broadcast.

[0066] After understanding the software and hardware architecture of this application, the following will be combined with... Figure 4This application describes an artificial intelligence-based news broadcasting method in its embodiments. Figure 4 This is a flowchart illustrating an artificial intelligence-based news broadcasting method provided in an embodiment of this application, specifically including the following steps:

[0067] Step S401: Obtain the text and chart information corresponding to the target news.

[0068] Specifically, relevant textual and graphical information can be extracted through the data source interface corresponding to the target news (such as the official interface of the news platform or the media content management system), or by parsing the presentation carrier of the target news (such as news webpage links, news document files, etc.). The textual information includes all plain text content such as the news title, lead, body paragraphs, key data interpretations, and event background descriptions, ensuring a complete understanding of the news's narrative logic and core information points. Graphical information includes, but is not limited to, the visualization elements embedded in the target news (such as data tables, trend charts, bar charts, infographics, etc.), and their accompanying chart titles, axis labels, data annotations, and other textual descriptions; no specific limitations are imposed here.

[0069] Step S402: Process the text information and the chart information to obtain the target broadcast text and its corresponding target feature information.

[0070] The specific steps of processing the text information and the chart information to obtain the target broadcast text and its corresponding target feature information include:

[0071] A1. Perform word segmentation on the text information to obtain multiple keywords;

[0072] A2. Perform correlation analysis on the multiple keywords to obtain the first text;

[0073] A3. The chart information is processed to obtain the second text;

[0074] A4. Determine the text content associated with the second text in the first text to obtain the third text;

[0075] A5. Determine the target similarity between the second text and the third text;

[0076] A6. If the target similarity is greater than the preset similarity threshold, the second text and the third text are fused to obtain the fourth text;

[0077] A7. Modify the first text according to the fourth text to obtain the target broadcast text;

[0078] A8. Extract features from the target broadcast text to obtain the target feature information.

[0079] In a specific embodiment, first, the text information (such as the title, lead, and body) of the target news can be disassembled into semantically independent lexical units through a preset word segmentation algorithm to obtain multiple keywords. Stop words without practical meaning (such as "of", "already", "in") are filtered, and core words are retained. Then, based on the semantic logic of the keywords (such as the association relationships between time, subject, action, and data), through techniques such as syntactic analysis and semantic role labeling, the scattered multiple keywords are recombined into a first text that conforms to the spoken language expression habit. Next, through a preset image recognition and data parsing technique, the core content in the chart information is extracted and converted into a second text. For example, if the chart information is a data chart (such as a line chart or a bar chart), the chart title, axis labels, data nodes, and trend changes of the data chart can be recognized and integrated to obtain the corresponding second text.

[0080] Then, through keyword matching (such as the coincidence degree of time, data indicators, and core subjects), the text content in the first text that has a logical association with the second text is located and used as the third text. Then, through a preset text similarity calculation algorithm (such as cosine similarity), the similarity degree between the second text and the third text is compared at the semantic level to obtain the target similarity. The higher the target similarity, the more consistent the description of the same content by the text information and the chart information, and the more suitable for fusion; otherwise, it is necessary to check whether there is information deviation to obtain accurate news content again.

[0081] If the target similarity is greater than the preset similarity threshold, the second text and the third text are fused to obtain a fourth text. In the process of fusion processing, the basic logic of the third text can be retained, and the chart information not covered by the text in the second text can be supplemented to form a more complete fourth text. The similarity threshold can be adjusted according to the broadcast scenario. For example, if the broadcast scenario is a government affairs broadcast or a professional financial interpretation, the similarity threshold can be increased to strictly screen the graphic and text combinations without ambiguity; if the broadcast scenario is a popular leisure broadcast, the similarity threshold can be decreased to prioritize ensuring the content richness and reduce the omission of effective information, which is not specifically limited here.

[0082] Finally, with the fourth text as the core, the fragments in the first text that are repeated with the fourth text are replaced, and the new information in the fourth text is supplemented to obtain the target broadcast text. By extracting features from the target broadcast text, the corresponding target feature information is obtained.

[0083] It is evident that by accurately splitting the keywords of textual information and translating the chart information into corresponding text, and then merging similar texts to generate the target broadcast text, the integration of news text and image information can be achieved. At the same time, the core features of the text can be accurately captured, providing a reliable basis for subsequent personalized broadcasts. This ensures both the integrity and accuracy of information transmission and improves the adaptability of the broadcast content.

[0084] Step S403: Analyze the target feature information to obtain reference content attributes and reference broadcast sentiment.

[0085] The specific steps for analyzing the target feature information to obtain reference content attributes and reference broadcast sentiment include:

[0086] B1. Extract multiple domain keywords, multiple time keywords, and multiple sentiment keywords from the target feature information;

[0087] B2. Determine the multiple domain categories corresponding to the multiple domain keywords according to the preset domain classification system;

[0088] B3. Determine the domain category that appears most frequently among the multiple domain categories as the target domain category;

[0089] B4. Determine the core event corresponding to the target broadcast text;

[0090] B5. Obtain at least one time keyword corresponding to the core event among the multiple time keywords;

[0091] B6. Obtain the target time corresponding to the time keyword that is closest to the current time among the at least one time keyword;

[0092] B7. Determine the timeliness of the target based on the time difference between the target time and the current time;

[0093] B8. Determine the reference content attributes based on the target domain category and the target timeliness;

[0094] B9. Determine the reference broadcast sentiment based on the multiple sentiment keywords.

[0095] In a specific embodiment, firstly, multiple domain keywords, multiple time keywords, and multiple sentiment keywords are extracted from the target feature information. Domain keywords represent keywords related to the field to which the news belongs, such as "stocks" and "market" in the financial field, "quantum computer" and "detector" in the technology field, and "disaster" and "rescue" in the social field. Time keywords represent time-related expressions associated with the news event, such as "yesterday," "next week," and "date." Sentiment keywords represent words with a clear emotional bias, such as positive terms like "success," "breakthrough," and "growth," negative terms like "accident," "loss," and "disaster," and neutral terms like "release," "statistics," and "notification," without specific limitations.

[0096] Then, based on a pre-defined domain classification system, the domain category corresponding to each of the multiple domain keywords is determined, resulting in multiple domain categories. For example, in this domain classification system, if the domain keyword is "Chang'e 6" or "lunar probe," then its corresponding domain category is "science and technology." Next, the frequency of occurrence of all domain keyword categories is counted, and the domain category with the highest frequency is selected as the target domain category.

[0097] Next, by analyzing the logical thread of the target broadcast text, its corresponding core event is identified. Then, from multiple time keywords, times directly related to this core event are filtered (excluding background times unrelated to the core event), resulting in at least one time keyword to ensure that the time analysis focuses on the core event. Then, the target time corresponding to the time keyword closest to the current time among at least one time keyword is obtained to ensure that the determination of the target's timeliness is based on the latest core event time.

[0098] Then, based on the time difference between the target time and the current time, the target timeliness is determined. The shorter the time difference, the higher the target timeliness. This timeliness can include any of the following: low timeliness, medium timeliness, and high timeliness, without specific limitations. Next, the target domain category and target timeliness are integrated to obtain reference content attributes. Finally, multiple sentiment keywords are analyzed to obtain reference broadcast sentiment.

[0099] It is evident that by accurately extracting and filtering keywords related to the field, time, and sentiment, and combining these with preset rules to determine the target field category, timeliness, and reference broadcast sentiment of the news, a reference content attribute that combines field attributes and timeliness is ultimately formed. This provides a precise and structured core basis for subsequent matching and adaptation of the target broadcast style, ensuring that the broadcast style is highly consistent with the essence of the news content.

[0100] For easier understanding, please refer to Figure 5 , Figure 5This is a flowchart illustrating a method for determining a reference broadcast sentiment according to an embodiment of this application. The specific steps for determining the reference broadcast sentiment based on the plurality of sentiment keywords include:

[0101] C1. Each of the multiple sentiment keywords is labeled to obtain multiple sentiment tendencies; the sentiment tendencies include any of the following: positive tendency, neutral tendency, and negative tendency.

[0102] C2. Determine multiple emotional scores corresponding to the multiple emotional tendencies according to the preset emotional scoring rules;

[0103] C3. Calculate the average of the multiple sentiment scores to obtain the average sentiment score;

[0104] C4. If the average emotion score is greater than or equal to the preset first emotion score threshold, then the reference broadcast emotion is determined to be a positive emotion.

[0105] C5. If the average emotion score is greater than the preset second emotion score threshold and less than the first emotion score threshold, then the reference broadcast emotion is determined to be a neutral emotion.

[0106] C6. If the average sentiment score is less than or equal to the second sentiment score threshold, then the reference broadcast sentiment is determined to be negative sentiment.

[0107] In a specific embodiment, firstly, each of the multiple emotional keywords can be labeled using a pre-set emotional lexicon or emotional classification model to obtain multiple emotional tendencies, which include any of the following: positive tendency, neutral tendency, and negative tendency. For example, emotional keywords such as "breakthrough" and "winning the championship" are directly labeled as "positive tendency" based on their semantics; emotional keywords such as "accident" and "decline" are labeled as "negative tendency"; and emotional keywords such as "statistics" and "release" that have no obvious emotional color are labeled as "neutral tendency".

[0108] Then, based on preset sentiment scoring rules, multiple sentiment scores are determined for multiple sentiment tendencies. For example, in this sentiment scoring rule, the sentiment score corresponding to "positive tendency" can be "+1", the sentiment score corresponding to "neutral tendency" can be "0", and the sentiment score corresponding to "negative tendency" can be "-1", without specific limitations. The average of the multiple sentiment scores is then calculated to obtain the average sentiment score.

[0109] Specifically, if the average emotion score is greater than or equal to a preset first emotion score threshold, the reference broadcast emotion is determined to be positive; if the average emotion score is greater than a preset second emotion score threshold but less than the first emotion score threshold, the reference broadcast emotion is determined to be neutral; and if the average emotion score is less than or equal to the second emotion score threshold, the reference broadcast emotion is determined to be negative.

[0110] It should be noted that the first and second sentiment scoring thresholds can be set to "+0.3" and "-0.3" respectively, and can also be flexibly adjusted according to the target domain category, without specific limitations here. For example, if the target domain category is finance, the first sentiment scoring threshold can be set to "0.4" and the second sentiment scoring threshold to "-0.4" to broaden the judgment range of neutral sentiment. If the target domain category is disaster, the second sentiment scoring threshold can be set to "-0.2" to more sensitively capture negative sentiment, thus adapting to different scenarios.

[0111] It is evident that by labeling, classifying, quantifying, and scoring emotional keywords, calculating the mean, and combining this with threshold judgment, vague emotional keywords can be transformed into clear reference emotional broadcasts. This approach is objective, accurate, and adaptable to multiple scenarios, providing a reliable basis for matching the tone of the broadcast emotional message.

[0112] Step S404: Determine the target broadcast style based on the reference content attributes and the reference broadcast sentiment.

[0113] For easier understanding, please refer to Figure 6 , Figure 6 This is a flowchart illustrating a method for determining a target broadcast style according to an embodiment of this application. The specific steps for determining the target broadcast style based on the reference content attributes and the reference broadcast sentiment include:

[0114] D1. Determine the reference broadcast tone corresponding to the target domain category based on the preset mapping relationship between domain categories and broadcast tone;

[0115] D2. Determine the target importance level corresponding to the target timeliness;

[0116] D3. Adjust the reference broadcast tone according to the target importance level to obtain the target broadcast tone;

[0117] D4. Determine the target broadcast sentiment based on the target broadcast tone and the reference broadcast sentiment;

[0118] D5. Determine the first broadcast tone and first broadcast speed corresponding to the target broadcast emotion;

[0119] D6. Determine the target broadcasting style based on the target broadcasting tone, the first broadcasting tone, and the first broadcasting speed.

[0120] In a specific embodiment, firstly, based on the preset mapping relationship between field categories and broadcast tone, the reference broadcast tone corresponding to the target field category is determined. For example, if the target field category is "finance," then its corresponding reference broadcast tone is "rigorous and solemn"; if the target field category is "technology," then its corresponding reference broadcast tone is "clear and objective"; if the target field category is "entertainment," then its corresponding reference broadcast tone is "relaxed and lively."

[0121] Next, determine the target importance level corresponding to the target's timeliness, where higher timeliness corresponds to higher importance. Then, adjust the intensity of the reference broadcast tone based on the target importance level to obtain the target broadcast tone.

[0122] Next, the target broadcast emotion is determined based on the target broadcast tone and the reference broadcast emotion. For example, if the target broadcast tone is "rigorous and solemn," and the reference broadcast emotion is "positive emotion," then the target broadcast emotion can be "rigorous and solemn positive emotion." Then, the first broadcast tone and first broadcast speed corresponding to the target broadcast emotion are determined. Each broadcast emotion corresponds to one broadcast tone and one broadcast speed, which is not specifically limited here. Finally, the target broadcast style is determined based on the target broadcast tone, the first broadcast tone, and the first broadcast speed. For example, if the target broadcast tone is "solid and composed," the first broadcast tone is "smooth ending," and the first broadcast speed is "medium-slow speed," the target broadcast style can be "a solid and composed style with a steady tone and medium-slow speed," which is not specifically limited here.

[0123] It is evident that by integrating the target broadcasting style from multiple perspectives, making it both consistent with the characteristics and importance of the news field and matching the emotional inclination of the target news, a deep fit between the broadcasting style and the essence of the news content is achieved, avoiding the problems of style uniformity or disconnection from the content.

[0124] Step S405: Obtain the target broadcasting requirements of the target users.

[0125] Specifically, the target user's desired broadcasting request can be obtained by responding to their input. This desired broadcasting request includes, but is not limited to, the target speech category and preferred broadcasting speed, which are not specifically limited here.

[0126] Step S406: Using a preset AI broadcasting model, the target broadcasting text is synthesized into speech based on the target broadcasting style and the target broadcasting requirements to obtain the target broadcasting speech.

[0127] The step of synthesizing the target broadcast text into speech using a preset AI broadcast model based on the target broadcast style and the target broadcast requirements to obtain the target broadcast speech includes the following steps:

[0128] E1. Determine the target speech category and speech rate preference corresponding to the target broadcasting requirement;

[0129] E2. Determine the reference broadcast timbre corresponding to the target speech category;

[0130] E3. Adjust the reference broadcast timbre according to the target broadcast tone to obtain the target broadcast timbre;

[0131] E4. If the first broadcast speed does not meet the broadcast speed preference, the first broadcast speed is adjusted according to the broadcast speed preference to obtain the target broadcast speed.

[0132] E5. If the first broadcast speed satisfies the broadcast speed preference, then the first broadcast speed is determined to be the target broadcast speed.

[0133] E6. Adjust the first broadcast tone according to the target broadcast tone and the target broadcast speed to obtain the target broadcast tone;

[0134] E7. Using the AI ​​broadcasting model, the target broadcasting text is synthesized into speech based on the target broadcasting timbre, the target broadcasting tone, the target broadcasting intonation, and the target broadcasting speed to obtain the target broadcasting speech.

[0135] In a specific embodiment, firstly, the target speech category and the speech rate preference are extracted from the target broadcasting requirements. The target speech category represents the speech type preferred by the target user (such as "young female voice", "elderly male voice", "children's voice", etc.), and the speech rate preference represents the speech rate range expected by the target user (such as "slow", "medium", "fast" or a specific numerical range).

[0136] Then, a corresponding reference broadcast timbre is matched based on the target speech category. For example, the reference broadcast timbre for "young female voice" could be "a female voice with a higher pitch and a clear timbre." Next, the reference broadcast timbre is adjusted according to the target broadcast tone to obtain the target broadcast timbre. For example, if the target broadcast tone is "rigorous and solemn," and the reference broadcast timbre is "a female voice with a higher pitch and a clear timbre," then its pitch can be lowered to obtain the target broadcast timbre, which will then be "a female voice with a mid-range pitch and a clear timbre," without further specific limitations.

[0137] Next, if the initial broadcast speed does not meet the preferred broadcast speed, it is adjusted according to the preference to obtain the target broadcast speed. If the initial broadcast speed meets the preference, it is determined as the target broadcast speed. Then, the initial broadcast tone is adjusted according to the target timbre and target broadcast speed to obtain the target broadcast tone. For example, if the target timbre is deep and resonant, and the target broadcast speed is fast, the intonation fluctuations of the initial broadcast tone can be reduced to obtain a smoother target broadcast tone.

[0138] Finally, the target broadcast timbre, target broadcast tone, target broadcast intonation, and target broadcast speed are input into the AI ​​broadcast model, and then the target broadcast text is synthesized to output the target broadcast speech.

[0139] It is evident that by coordinating and adjusting multiple parameters, the speech synthesis not only meets the user's personalized preferences but also ensures that the speech style matches the news content, ultimately generating a broadcast voice that is both personalized and content-adaptive, effectively improving the quality of news broadcasting and the user's listening experience.

[0140] Step S407: In response to the voice broadcast operation of the target user, broadcast the target broadcast voice.

[0141] Specifically, when a target user triggers a voice broadcast operation (such as clicking the "play" button, saying the voice command "play this news", etc.), the system will immediately retrieve the generated target broadcast voice file or trigger the AI ​​broadcast model in real time to generate the target broadcast voice for broadcast.

[0142] In one possible embodiment, feedback data from the target user regarding the target broadcast voice can be obtained; key information, optimization direction, and optimization intensity corresponding to the feedback data can be determined; broadcast parameters to be optimized can be determined based on the key information; the broadcast parameters to be optimized include at least one of the following: the target broadcast timbre, the target broadcast tone, the target broadcast intonation, and the target broadcast speed; the broadcast parameters to be optimized can be optimized based on the optimization direction and the optimization intensity to obtain optimized target broadcast parameters; the target broadcast parameters are used to adjust and regenerate the target broadcast voice.

[0143] Specifically, a feedback window can be provided for target users to input text feedback, thereby obtaining feedback data on the target broadcast voice. This feedback data is then analyzed to determine the corresponding key information, optimization direction, and optimization intensity. For example, if the feedback data is "The speech speed is too fast, I can't hear the news content clearly," then the key information is "The speech speed is too fast," the optimization direction is "reduce the speech speed," and the optimization intensity is "significant adjustment." Next, based on the key information, the broadcast parameters to be optimized are determined. These parameters include at least one of the following: target broadcast timbre, target broadcast tone, target broadcast intonation, and target broadcast speed. For example, if the key information is "The speech speed is too fast," then the broadcast parameter to be optimized is the target broadcast speed. Finally, the broadcast parameters to be optimized are optimized according to the optimization direction and optimization intensity to obtain the optimized target broadcast parameters. These optimized parameters are used to adjust and regenerate the target broadcast voice.

[0144] It is evident that by collecting feedback data from target users and optimizing the parameters of the target broadcast voice, the target broadcast voice can continuously match user preferences, thereby improving the quality of news broadcasting.

[0145] For easier understanding, please refer to Figure 7 , Figure 7 This is a schematic diagram of a feedback optimization process provided in an embodiment of this application. The process begins by initiating a feedback optimization process and then collecting feedback data from target users on the target broadcast voice. Next, parameter optimization is performed based on the feedback data, adjusting and optimizing relevant parameters of the target broadcast voice (such as speech rate, tone, and timbre). Finally, the broadcast voice, i.e., the target broadcast voice, is regenerated based on the optimized parameters, and the current feedback optimization process ends. The process can be restarted again based on new feedback data in subsequent iterations.

[0146] The above primarily describes the solutions of the embodiments of this application from the perspective of the method execution process. It is understood that, in order to achieve the above functions, the electronic device includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments provided herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0147] This application embodiment can divide the electronic device into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0148] When dividing each function into modules according to its corresponding function. Figure 8 This is a functional module block diagram of an artificial intelligence-based news broadcasting device 800 provided in an embodiment of this application. The artificial intelligence-based news broadcasting device 800 includes an acquisition module 810, a processing module 820, an analysis module 830, a determination module 840, a synthesis module 850, and a broadcasting module 860, wherein:

[0149] The acquisition module 810 is used to acquire the text information and chart information corresponding to the target news;

[0150] The processing module 820 is used to process the text information and the chart information to obtain the target broadcast text and its corresponding target feature information.

[0151] The analysis module 830 is used to analyze the target feature information to obtain reference content attributes and reference broadcast sentiment.

[0152] The determining module 840 is used to determine the target broadcasting style based on the reference content attributes and the reference broadcasting sentiment.

[0153] The acquisition module 810 is also used to acquire the target user's target broadcasting requirements;

[0154] The synthesis module 850 is used to synthesize the target broadcast text into speech based on the target broadcast style and the target broadcast requirements using a preset AI broadcast model, so as to obtain the target broadcast speech.

[0155] The broadcast module 860 is used to respond to the voice broadcast operation of the target user and broadcast the target voice.

[0156] Optionally, in processing the text information and the chart information to obtain the target broadcast text and its corresponding target feature information, the processing module 820 is specifically used for:

[0157] The text information is segmented to obtain multiple keywords;

[0158] A correlation analysis was performed on the multiple keywords to obtain the first text;

[0159] The chart information is processed to obtain the second text;

[0160] Determine the text content associated with the second text in the first text to obtain the third text;

[0161] Determine the target similarity between the second text and the third text;

[0162] If the target similarity is greater than a preset similarity threshold, the second text and the third text will be fused to obtain the fourth text;

[0163] The first text is modified according to the fourth text to obtain the target broadcast text;

[0164] Feature extraction is performed on the target broadcast text to obtain the target feature information.

[0165] Optionally, in analyzing the target feature information to obtain reference content attributes and reference broadcast sentiment, the analysis module 830 is specifically used for:

[0166] Extract multiple domain keywords, multiple time keywords, and multiple sentiment keywords from the target feature information;

[0167] The multiple domain categories corresponding to the multiple domain keywords are determined according to the preset domain classification system;

[0168] The domain category that appears most frequently among the multiple domain categories is identified as the target domain category.

[0169] Determine the core event corresponding to the target broadcast text;

[0170] Obtain at least one time keyword corresponding to the core event from among the multiple time keywords;

[0171] Obtain the target time corresponding to the time keyword closest to the current time among the at least one time keyword;

[0172] The timeliness of the target is determined based on the time difference between the target time and the current time;

[0173] The reference content attributes are determined based on the target domain category and the target timeliness.

[0174] The reference broadcast sentiment is determined based on the multiple sentiment keywords.

[0175] Optionally, in determining the reference broadcast sentiment based on the plurality of sentiment keywords, the analysis module 830 is further specifically used for:

[0176] Each of the multiple sentiment keywords is labeled to obtain multiple sentiment tendencies; the sentiment tendencies include any one of the following: positive tendency, neutral tendency, and negative tendency.

[0177] Multiple emotional scores corresponding to the multiple emotional tendencies are determined according to preset emotional scoring rules;

[0178] Calculate the average of the multiple sentiment scores to obtain the average sentiment score;

[0179] If the average sentiment score is greater than or equal to a preset first sentiment score threshold, then the reference broadcast sentiment is determined to be a positive sentiment.

[0180] If the average sentiment score is greater than a preset second sentiment score threshold and less than the first sentiment score threshold, then the reference broadcast sentiment is determined to be neutral.

[0181] If the average sentiment score is less than or equal to the second sentiment score threshold, then the reference broadcast sentiment is determined to be negative.

[0182] Optionally, in determining the target broadcast style based on the reference content attributes and the reference broadcast sentiment, the determining module 840 is specifically used for:

[0183] Based on the preset mapping relationship between domain categories and broadcast tone, the reference broadcast tone corresponding to the target domain category is determined;

[0184] Determine the target importance level corresponding to the target timeliness;

[0185] The reference broadcast tone is adjusted according to the target importance level to obtain the target broadcast tone;

[0186] The target broadcast sentiment is determined based on the target broadcast tone and the reference broadcast sentiment.

[0187] Determine the first broadcast tone and first broadcast speed corresponding to the target broadcast emotion;

[0188] The target broadcast style is determined based on the target broadcast tone, the first broadcast intonation, and the first broadcast speed.

[0189] Optionally, in the step of synthesizing speech from the target broadcast text using a preset AI broadcast model based on the target broadcast style and the target broadcast requirements to obtain the target broadcast speech, the synthesis module 850 is specifically used for:

[0190] Determine the target speech category and speech rate preference corresponding to the target broadcasting requirement;

[0191] Determine the reference broadcast timbre corresponding to the target speech category;

[0192] The reference broadcast timbre is adjusted according to the target broadcast tone to obtain the target broadcast timbre;

[0193] If the first broadcast speed does not meet the broadcast speed preference, the first broadcast speed is adjusted according to the broadcast speed preference to obtain the target broadcast speed.

[0194] If the first broadcast speed satisfies the broadcast speed preference, then the first broadcast speed is determined to be the target broadcast speed;

[0195] The first broadcast tone is adjusted according to the target broadcast tone and the target broadcast speed to obtain the target broadcast tone;

[0196] The AI ​​broadcasting model synthesizes the target broadcasting text into speech based on the target broadcasting timbre, the target broadcasting tone, the target broadcasting intonation, and the target broadcasting speed to obtain the target broadcasting speech.

[0197] Optionally, the synthesis module 850 is further specifically used for:

[0198] Obtain feedback data from the target user regarding the target broadcast voice;

[0199] Determine the key information, optimization direction, and optimization intensity corresponding to the feedback data;

[0200] The broadcast parameters to be optimized are determined based on the key information; the broadcast parameters to be optimized include at least one of the following: the target broadcast timbre, the target broadcast tone, the target broadcast intonation, and the target broadcast speed;

[0201] The broadcast parameters to be optimized are optimized according to the optimization direction and the optimization intensity to obtain the optimized target broadcast parameters; the target broadcast parameters are used to adjust and regenerate the target broadcast voice.

[0202] It is evident that by processing the text and image information of news, analyzing its content attributes and broadcasting sentiment to determine the broadcasting style, and then combining it with users' personalized needs, the AI ​​broadcasting model synthesizes adapted voice and responds to user operations to complete the broadcast, thereby improving the quality of news broadcasting.

[0203] It should be noted that the specific implementation of each operation can be described in the corresponding description of the method embodiments shown above. The artificial intelligence-based news broadcasting device 800 can be used to execute the above method embodiments of this application, and will not be described again here.

[0204] This application also provides a computer-readable storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments, wherein the computer includes an electronic device.

[0205] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may include an electronic device.

[0206] It should be noted that, for the sake of simplicity, the above embodiments are all described as a series of actions. Those skilled in the art should understand that this application is not limited to the described order of actions, as some steps in the embodiments of this application can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions, steps, modules, or units involved are not necessarily essential to the embodiments of this application.

[0207] In the above embodiments, the descriptions of each embodiment in this application have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0208] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

[0209] The steps of the methods or algorithms described in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in RAM, flash memory, ROM, EPROM, electrically erasable programmable read-only memory (EEPROM), registers, hard disk, portable hard disk, read-only optical disk (CD-ROM), or any other form of storage medium well known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Furthermore, the ASIC can reside in a terminal device or management device. Alternatively, the processor and storage medium can exist as discrete components in the terminal device or management device.

[0210] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in the embodiments of this application can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0211] The modules / units included in the various devices and products described in the above embodiments can be software modules / units, hardware modules / units, or a combination of both. For example, for devices and products applied to or integrated into a chip, all modules / units can be implemented using hardware methods such as circuits, or at least some modules / units can be implemented using software programs that run on a processor integrated within the chip, while the remaining (if any) modules / units can be implemented using hardware methods such as circuits. For devices and products applied to or integrated into a chip module, all modules / units can be implemented using hardware methods such as circuits. Different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or different components of the chip module, or at least some modules / units can be implemented using hardware methods such as circuits. The implementation is achieved through a software program that runs on the processor integrated within the chip module. The remaining modules / units (if any) can be implemented using hardware methods such as circuits. For various devices and products applied to or integrated into terminal equipment, each of their modules / units can be implemented using hardware methods such as circuits. Different modules / units can be located in the same component (e.g., chip, circuit module, etc.) or different components within the terminal equipment. Alternatively, at least some modules / units can be implemented through a software program that runs on the processor integrated within the terminal equipment, while the remaining modules / units (if any) can be implemented using hardware methods such as circuits.

[0212] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the embodiments of this application. It should be understood that the above descriptions are merely specific embodiments of the embodiments of this application and are not intended to limit the protection scope of the embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solutions of the embodiments of this application should be included within the protection scope of the embodiments of this application.

Claims

1. A news broadcasting method based on artificial intelligence, characterized in that, The method includes: Retrieve the text and chart information corresponding to the target news; The text information and the chart information are processed to obtain the target broadcast text and its corresponding target feature information; The target feature information is analyzed to obtain reference content attributes and reference broadcast sentiment. The target broadcast style is determined based on the reference content attributes and the reference broadcast sentiment. Obtain the target users' desired broadcasting needs; Using a preset AI broadcasting model, the target broadcasting text is synthesized into speech based on the target broadcasting style and the target broadcasting requirements to obtain the target broadcasting speech; In response to the voice broadcast operation of the target user, broadcast the target broadcast voice. The step of processing the text information and the chart information to obtain the target broadcast text and its corresponding target feature information includes: The text information is segmented to obtain multiple keywords; A correlation analysis was performed on the multiple keywords to obtain the first text; The chart information is processed to obtain the second text; Determine the text content associated with the second text in the first text to obtain the third text; Determine the target similarity between the second text and the third text; If the target similarity is greater than a preset similarity threshold, the second text and the third text will be fused to obtain the fourth text; The first text is modified according to the fourth text to obtain the target broadcast text; Feature extraction is performed on the target broadcast text to obtain the target feature information; The analysis of the target feature information to obtain reference content attributes and reference broadcast sentiment includes: Extract multiple domain keywords, multiple time keywords, and multiple sentiment keywords from the target feature information; The multiple domain categories corresponding to the multiple domain keywords are determined according to the preset domain classification system; The domain category that appears most frequently among the multiple domain categories is identified as the target domain category. Determine the core event corresponding to the target broadcast text; Obtain at least one time keyword corresponding to the core event from among the multiple time keywords; Obtain the target time corresponding to the time keyword closest to the current time among the at least one time keyword; The timeliness of the target is determined based on the time difference between the target time and the current time; The reference content attributes are determined based on the target domain category and the target timeliness. The reference broadcast sentiment is determined based on the multiple sentiment keywords.

2. The method as described in claim 1, characterized in that, Determining the reference broadcast sentiment based on the plurality of sentiment keywords includes: Each of the multiple sentiment keywords is labeled to obtain multiple sentiment tendencies; the sentiment tendencies include any one of the following: positive tendency, neutral tendency, and negative tendency. Multiple emotional scores corresponding to the multiple emotional tendencies are determined according to preset emotional scoring rules; Calculate the average of the multiple sentiment scores to obtain the average sentiment score; If the average sentiment score is greater than or equal to a preset first sentiment score threshold, then the reference broadcast sentiment is determined to be a positive sentiment. If the average emotion score is greater than a preset second emotion score threshold and less than the first emotion score threshold, then the reference broadcast emotion is determined to be a neutral emotion. If the average sentiment score is less than or equal to the second sentiment score threshold, then the reference broadcast sentiment is determined to be negative.

3. The method as described in claim 2, characterized in that, The step of determining the target broadcasting style based on the reference content attributes and the reference broadcasting sentiment includes: Based on the preset mapping relationship between domain categories and broadcast tone, the reference broadcast tone corresponding to the target domain category is determined; Determine the target importance level corresponding to the target timeliness; The reference broadcast tone is adjusted according to the target importance level to obtain the target broadcast tone; The target broadcast sentiment is determined based on the target broadcast tone and the reference broadcast sentiment. Determine the first broadcast tone and first broadcast speed corresponding to the target broadcast emotion; The target broadcast style is determined based on the target broadcast tone, the first broadcast intonation, and the first broadcast speed.

4. The method as described in claim 3, characterized in that, The step of synthesizing the target broadcast text into speech using a preset AI broadcast model, based on the target broadcast style and the target broadcast requirements, to obtain the target broadcast speech includes: Determine the target speech category and speech rate preference corresponding to the target broadcasting requirement; Determine the reference broadcast timbre corresponding to the target speech category; The reference broadcast timbre is adjusted according to the target broadcast tone to obtain the target broadcast timbre; If the first broadcast speed does not meet the broadcast speed preference, the first broadcast speed is adjusted according to the broadcast speed preference to obtain the target broadcast speed. If the first broadcast speed satisfies the broadcast speed preference, then the first broadcast speed is determined to be the target broadcast speed; The first broadcast tone is adjusted according to the target broadcast tone and the target broadcast speed to obtain the target broadcast tone; The AI ​​broadcasting model synthesizes the target broadcasting text into speech based on the target broadcasting timbre, the target broadcasting tone, the target broadcasting intonation, and the target broadcasting speed to obtain the target broadcasting speech.

5. The method as described in claim 4, characterized in that, The method further includes: Obtain feedback data from the target user regarding the target broadcast voice; Determine the key information, optimization direction, and optimization intensity corresponding to the feedback data; The broadcast parameters to be optimized are determined based on the key information; the broadcast parameters to be optimized include at least one of the following: the target broadcast timbre, the target broadcast tone, the target broadcast intonation, and the target broadcast speed; The broadcast parameters to be optimized are optimized according to the optimization direction and the optimization intensity to obtain the optimized target broadcast parameters; the target broadcast parameters are used to adjust and regenerate the target broadcast voice.

6. An artificial intelligence-based news broadcasting device, used to perform the method as described in any one of claims 1-5, characterized in that, The device includes an acquisition module, a processing module, an analysis module, a determination module, a synthesis module, and a broadcasting module, wherein: The acquisition module is used to acquire the text information and chart information corresponding to the target news; The processing module is used to process the text information and the chart information to obtain the target broadcast text and its corresponding target feature information; The analysis module is used to analyze the target feature information to obtain reference content attributes and reference broadcast sentiment. The determining module is used to determine the target broadcasting style based on the reference content attributes and the reference broadcasting sentiment; The acquisition module is also used to acquire the target user's target broadcasting requirements; The synthesis module is used to synthesize the target broadcast text into speech based on the target broadcast style and the target broadcast requirements using a preset AI broadcast model, so as to obtain the target broadcast speech. The broadcast module is used to respond to the voice broadcast operation of the target user and broadcast the target voice.

7. An electronic device, characterized in that, include: Processor, memory, communication interface, and one or more programs; The one or more programs are stored in the memory and configured to be executed by the processor, the programs including instructions for performing the steps of the method as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Multi-modal data fusion method based on position sensitive optimization

    CN119339193A

  • Intelligent news broadcasting method and system based on artificial intelligence

    CN119441573A

  • Voice generation method and device based on multi-modal fusion, equipment and medium

    CN120048243A