SSML editing method and device

Through the visual editing interface based on the DOM structure, SSML tag editing is simplified, and the complex operation and compatibility problems of existing tools are solved, efficient SSML data generation and speech synthesis effect preview are realized, and multi-tone word recognition and word count control are supported.

CN120297237APending Publication Date: 2025-07-11SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510410014.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing SSML editing tools are complex in operation, resulting in high threshold for use, making it difficult to quickly generate SSML code that meets the needs, and compatibility problems occur frequently during cross-platform use, making speech synthesis effects in real time, cannot accurately identify and convert polyphonic text, and word count control is not fine.

Method used

It provides a visual editing interface based on the DOM structure. Through multiple tag controls and attribute adjustment windows, it supports tag insertion, deletion and adjustment of text content, realizes two-way conversion between SSML and HTML formats, and provides real-time speech synthesis effect preview and multi-sound word recognition functions to ensure that the generated SSML files meet the requirements of the speech synthesis engine.

Benefits of technology

Simplify SSML tag editing through the visual interface, improve voice editing efficiency, support cross-platform format compatibility, realize real-time voice synthesis effect preview and multi-tone word processing, ensuring that the generated SSML data meets the requirements of the voice synthesis engine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297237A_ABST
    Figure CN120297237A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an SSML editing method and related equipment / products, and relates to the technical field of multimedia. The SSML editing method comprises the steps that an editing interface based on a DOM structure is displayed, the editing interface is configured with a plurality of label controls corresponding to a plurality of label types, and each label type has a plurality of label attributes; receiving text content through the editing interface and determining a target label type and a target label attribute selected for the text content; according to the text content, the target label type and the target label attribute, creating a text node and a label node in a DOM structure; the text nodes are converted into SSML texts, and the label nodes are converted into SSML labels; and synthesizing the SSML data according to the SSML text and the SSML label. According to the technical scheme provided by the embodiment of the invention, the label is inserted and edited in the text content, so that the text content with the label can be efficiently and accurately converted into the SSML data for guiding speech synthesis, and the speech editing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of data processing, and in particular, to a method and device for SSML editing, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the rapid development of artificial intelligence and natural language processing technologies, speech synthesis technology has become an important part of fields such as human-computer interaction, intelligent assistants, and barrier-free services. In a speech synthesis system, SSML technology is widely used to control the performance of synthesized speech, such as pitch, speech rate, volume, pause, etc.

[0003] Currently, most SSML editing tools adopt a text-based editing method, and users need to manually write SSML tags. This method not only increases the complexity of operations but also results in a relatively high usage threshold.

[0004] It should be noted that the above content is not necessarily prior art and is not used to limit the patent protection scope of the present application. Summary of the Invention

[0005] Embodiments of the present application provide a method and device for SSML editing, a computer device, a computer-readable storage medium, and a computer program product to solve or alleviate one or more of the above technical problems.

[0006] One aspect of the embodiments of the present application provides an SSML editing method, and the method includes: Display an editing interface based on a DOM structure, where the editing interface is configured with a plurality of tag controls, the plurality of controls correspond to a plurality of tag types, and each tag type has a plurality of tag attributes; Receive text content through the editing interface and determine a target tag type and a target tag attribute selected for the text content; Create a text node and a tag node in the DOM structure according to the text content, the target tag type, and the target tag attribute; Convert the text node in the DOM structure into SSML text, and convert the tag node in the DOM structure into an SSML tag; Synthesize SSML data according to the SSML text and the SSML tag, and the SSML data is used to guide speech synthesis for the text content.

[0007] Optionally, the editing interface includes a text input area and a tag selection area, and the tag selection area includes the plurality of tag controls; receiving text content through the editing interface and determining a target tag type and a target tag attribute selected for the text content includes: Receiving the text content through the text input area, where the text content includes multiple text paragraphs; In response to selecting any one of the text paragraphs and a target tag type selected from the multiple tag types for the text paragraph, creating a tag corresponding to the target tag type after the text paragraph; or in response to selecting any two text paragraphs and a target tag type selected from the multiple tag types for the two text paragraphs, creating a tag corresponding to the target tag type between the two text paragraphs; In response to selecting the tag, presenting a tag attribute adjustment window, where the tag attribute adjustment window includes multiple tag attributes or tag attribute input boxes corresponding to the target tag type; In response to selecting one of the tag attributes, determining the selected tag attribute as the target tag attribute, or receiving the target tag attribute through the tag attribute input box.

[0008] Optionally, creating text nodes and tag nodes in the DOM structure according to the text content, the target tag type, and the target tag attribute, including: Creating corresponding text nodes in the DOM structure according to the text paragraphs; Creating corresponding tag nodes in the DOM structure according to the target tag type and the target tag attribute corresponding to each tag; Wherein, according to the positional relationship between the text paragraph and the tag, adding the tag node after the corresponding text node or inserting it between the corresponding two text nodes.

[0009] Optionally, converting the text nodes in the DOM structure into SSML text, and converting the tag nodes in the DOM structure into SSML tags, including: Creating an SSML container; Generating corresponding SSML text and SSML tags in sequence in the SSML container according to the sorting of the text nodes and the tag nodes in the DOM structure; Wherein, for each text node, extracting valid text from the corresponding text paragraph through a regular expression and converting the valid text into the SSML text; for each tag node, extracting the target tag type and the target tag attribute from the tag node and converting the target tag type and the target tag attribute into the SSML tag.

[0010] Optionally, the SSML editing method further includes: Generating corresponding historical editing records according to the SSML data; Display the historical editing records on the editing interface; In response to selecting the historical editing record, display the text content and the selected target tag type and target tag attributes for the text content on the editing interface.

[0011] Optionally, in response to selecting the historical editing record, display the text content and the selected target tag type and target tag attributes for the text content on the editing interface, including: In response to selecting the historical editing record, obtain the corresponding SSML data; Parse the SSML data through a pre-constructed DOM parser and convert it into a DOM structure; wherein, the SSML data includes SSML text and SSML tags, the DOM structure includes text nodes and tag nodes, the SSML text corresponds to the text nodes, and the SSML tags correspond to the tag nodes; According to the DOM structure, render the text content and the selected target tag type and target tag attributes for the text content on the editing interface.

[0012] Optionally, parse the SSML data through a pre-constructed DOM parser and convert it into a DOM structure, including: Extract the target tag type and target tag attributes from the SSML tags; Obtain multiple tag attributes corresponding to the target tag type, and create the tag nodes according to the target tag type, target tag attributes, and the multiple tag attributes; Correspondingly, according to the DOM structure, render the text content and the selected target tag type and target tag attributes for the text content on the editing interface, including: Render multiple text paragraphs and corresponding tags on the editing interface according to the text nodes and the tag nodes; Wherein, each tag carries the target tag attributes; when the tag is selected, a tag attribute adjustment window is displayed on the editing interface, and the tag attribute adjustment window includes the multiple tag attributes.

[0013] Another aspect of the embodiments of the present application provides an SSML editing device, and the device includes: A display module for displaying an editing interface based on a DOM structure, where the editing interface is configured with multiple tag controls, the multiple controls correspond to multiple tag types, and each tag type has multiple tag attributes; A receiving module for receiving text content through the editing interface and determining the selected target tag type and target tag attributes for the text content; A creation module, configured to create text nodes and tag nodes in the DOM structure according to the text content, the target tag type, and the target tag attributes. A conversion module, configured to convert the text nodes in the DOM structure into SSML text, and convert the tag nodes in the DOM structure into SSML tags. A synthesis module, configured to synthesize SSML data according to the SSML text and the SSML tags, where the SSML data is used to guide speech synthesis for the text content.

[0014] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0015] Another aspect of the embodiments of the present application provides a computer-readable storage medium, where computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed by a processor, the method as described above is implemented.

[0016] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0017] The embodiments of the present application adopting the above technical solutions may include the following advantages: By receiving text content and selected tag types and tag attributes through an editing interface based on the DOM structure. Then, according to the text content, tag types, and tag attributes, text nodes and tag nodes are created in the DOM structure, and the text nodes and tag nodes are converted into SSML text and SSML tags. Based on the SSML text and SSML tags, SSML data for guiding speech synthesis can be generated. The embodiments of the present application support inserting and editing tags in text content through a visual interface, and efficiently and accurately converting the tagged text content into SSML data for guiding speech synthesis, thereby improving the efficiency of speech editing. Description of the Drawings

[0018] The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary implementation manners of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0019] Figure 1 Schematically shows the operating environment diagram of the SSML editing method according to Embodiment 1 of the present application; Figure 2 Schematically shows the flowchart of the SSML editing method according to Embodiment 1 of the present application; Figure 3 Schematically shows Figure 2 The sub-step flowchart of step S202 in; Figure 4 Schematically shows Figure 2 The sub-step flowchart of step S204 in; Figure 5 Schematically shows Figure 2 The sub-step flowchart of step S206 in; Figure 6 Schematically shows the new flowchart of the SSML editing method according to Embodiment 1 of the present application; Figure 7 Schematically shows Figure 6 The sub-step flowchart of step S604 in; Figure 8A and Figure 8B Schematically shows the exemplary application flowchart; Figure 9 Schematically shows the editing interface display diagram of the embodiment of the present application; Figure 10 Schematically shows the component relationship diagram of the embodiment of the present application; Figure 11 Schematically shows the block diagram of the SSML editing device according to Embodiment 2 of the present application; and Figure 12 Schematically shows the hardware architecture diagram of the computer device according to Embodiment 3 of the present application. Detailed implementation manners

[0020] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0021] It should be noted that in the embodiments of the present application, the descriptions involving "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second" may explicitly or implicitly include at least one of such features. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or inability to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0022] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the sequence of execution of the steps, but are only used to facilitate the description of the present application and distinguish each step. Therefore, it should not be construed as a limitation to the present application.

[0023] First, the following provides the term explanations involved in the present application: XML (Extensible Markup Language): A general markup language used for storing and transmitting data, widely applied in fields such as data exchange, configuration files, and document structure definition.

[0024] SSML (Speech Synthesis Markup Language): A markup language based on XML, used to control the output of a speech synthesis system. By embedding tags and attributes, various aspects of the synthesized speech can be controlled. Among them, tags are used to define the behavior of speech synthesis. Attributes are used to provide additional information for the tags.

[0025] HTML (HyperText Markup Language): A markup language used to create web pages, defining the structure and content of web pages through tags and attributes. Among them, tags are used to define web page elements. Attributes are used to provide additional information for the tags.

[0026] DOM (Document Object Model): A tree-like structure representation of a web page, used to describe the logical structure of a document. In SSML editing, the DOM structure is used to store text content, tags, and their hierarchical relationships.

[0027] Vue framework: A progressive JavaScript framework for building user interfaces, supporting component-based development, and capable of splitting the interface into independent and reusable components. In addition, Vue provides a directive system that can be used to handle attribute binding and event binding of tags, simplifying the development of DOM operations and interaction logic.

[0028] Secondly, to facilitate the understanding of the technical solution provided by the embodiments of the present application by those skilled in the art, the related technologies are described below: The SSML editing tool lacks a graphical interface, making it difficult to quickly generate SSML code that meets requirements, and it does not support the efficient conversion between SSML and formats such as HTML. Frequent compatibility issues occur when used across platforms. During the editing process, users cannot preview the speech synthesis effect in real time, resulting in cumbersome debugging and low efficiency. For polyphonic text, the SSML editing tool cannot accurately identify and convert it. The control of the number of characters is not precise enough, easily causing the generated SSML to exceed the character limit.

[0029] For this reason, the embodiments of the present application provide an SSML editing technical solution. In this technical solution: (1) Simplify the editing of SSML tags through the visual interface of the SSML editor. For example, insert, delete, and adjust tags without manually writing complex code. (2) This SSML editor also supports the bidirectional conversion between SSML and HTML formats, effectively solving the format compatibility problem when used across platforms. (3) Provide a preview function for the real-time speech synthesis effect, significantly improving the editing efficiency. (4) Provide a character limit function and a polyphonic character recognition function, automatically identifying and processing polyphonic characters to ensure that the generated SSML file meets the requirements of the speech synthesis engine. (5) The generation and interaction of SSML tags can be dynamically controlled by binding HTML tags to the attributes and events of the Vue framework. See the following for details.

[0030] Finally, for the convenience of understanding, an exemplary operating environment is provided below.

[0031] As Figure 1 shown, the operating environment diagram includes: a service platform 2, clients (4A, 4B,..., 4N).

[0032] The service platform 2 can connect to the clients (4A, 4B,..., 4N) through a network.

[0033] The service platform 2 can be a single server, a server cluster, or a cloud computing service center.

[0034] The service platform 2 can provide speech synthesis services and the like to the clients.

[0035] The service platform 2 can be located in a data center such as a single location, or distributed in different geographical locations (for example, in multiple locations). The service platform 2 can provide services via a network. The network includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network can include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links, and their combinations, or wireless links, such as cellular links, satellite links, Wi-Fi links, etc.

[0036] Clients (4A, 4B, …, 4N) can be configured to access the content and services of service platform 2, such as text-to-speech services. Of course, both SSML editing and SSML running (text-to-speech) in the embodiments of the present application can be executed locally on the client. Clients (4A, 4B, …, 4N) can include electronic devices with or external to a display panel, such as mobile devices, tablet devices, laptop computers, workstations, virtual reality devices, gaming devices, digital streaming devices, vehicle terminals, smart TVs, set-top boxes, etc., and can also include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing device can load a virtual machine based on a virtual image and / or other data defining specific software (e.g., operating system, dedicated application, server) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices.

[0037] Clients (4A, 4B, …, 4N) can be associated with one or more users. A single user can also use one or more of clients (4A, 4B, …, 4N) to access service platform 2. Clients (4A, 4B, …, 4N) can travel to various locations and use different networks to access service platform 2.

[0038] Clients (4A, 4B, …, 4N) can include an interface. The interface can include a touchpad, a touch screen, a mouse, a keyboard, or other sensing elements. For example, the input element can be configured to receive user instructions, and the user instructions can cause clients (4A, 4B, …, 4N) to perform various operations, such as inputting text, selecting a label control, previewing speech, etc.

[0039] It should be noted that the above devices are exemplary, and in different scenarios or according to different requirements, the number and types of devices can be adjusted.

[0040] The technical solutions of the present application will be introduced through multiple embodiments below. It should be noted that these embodiments can be implemented in many different forms and should not be construed as being limited only to the embodiments described herein.

[0041] Embodiment 1 Figure 2 A flowchart of the SSML editing method according to Embodiment 1 of the present application is schematically shown.

[0042] As Figure 2 shown, the SSML editing method can include steps S200~S208, where: Step S200: Display an editing interface based on the DOM structure. The editing interface is configured with multiple label controls. The multiple controls correspond to multiple label types, and each label type has multiple label attributes. Step S202: Receive the text content through the editing interface and determine the target label type and target label attributes selected for the text content. Step S204: Create a text node and a label node in the DOM structure according to the text content, the target label type, and the target label attributes. Step S206: Convert the text node in the DOM structure into SSML text, and convert the label node in the DOM structure into an SSML label. Step S208: Synthesize SSML data according to the SSML text and the SSML label. The SSML data is used to guide the speech synthesis for the text content.

[0043] The SSML editing method provided in this embodiment receives the text content and the selected label type and label attributes through an editing interface based on the DOM structure. Then, according to the text content, the label type, and the label attributes, a text node and a label node are created in the DOM structure. And the text node and the label node are converted into SSML text and an SSML label, and finally the SSML data for speech synthesis is obtained. The embodiments of the present application can insert and edit labels in the text content through a visual interface, support efficiently and accurately converting the text content with labels into SSML data for guiding speech synthesis, thereby improving the speech editing efficiency.

[0044] The following Figure 2 elaborates in detail on each step in steps S200 to S208 and other optional steps.

[0045] Step S200 Display an editing interface based on the DOM structure. The editing interface is configured with multiple label controls. The multiple controls correspond to multiple label types, and each label type has multiple label attributes.

[0046] The editing interface is a visual tool for the user to interact with the DOM structure, and dynamically updates the DOM structure through interface operations (such as dragging and clicking). When inserting, deleting, or editing labels in the editing interface, the editing interface can update the DOM structure in real time. The changes in the DOM structure can be synchronously reflected in the editing interface, effectively ensuring the consistency between user operations and data.

[0047] The label control may be an interactive element configured in the editing interface, and each label control may correspond to a label type. The label control may include: a polyphonic character selection control, a number reading control, a pronunciation replacement control, a pause insertion control, etc. The label type may be used to control the performance of speech synthesis. The label attribute may be a parameter of the label type. For example, the user can determine the label type by selecting the label control. For example, by selecting the pause insertion control, the corresponding label type is "pause", etc. After determining the label type, the label attribute (such as the pause time) can be set for this label type in the editing interface. For example, pause for 0.5 seconds.

[0048] Step S202 , receive the text content through the editing interface and determine the target label type and target label attribute selected for the text content.

[0049] The user can select a text paragraph in various ways such as by mouse or touch, and then select the label type to be applied to this text paragraph and set the corresponding label attribute (i.e., the target label type and target label attribute). For example, the user selects the target label type from the toolbar or the drop-down menu. Example: Select the "pronunciation replacement" control, corresponding to <phoneme>Label type. Set label attributes in various ways such as through an input box or a drop-down menu. For example, set alphabet = "pinyin" and ph = "yín háng". In some embodiments, the target label type and target label attributes can be added to the selected text paragraph in various ways such as by clicking the "Apply" button.

[0050] In an alternative embodiment, the editing interface includes a text input area and a label selection area, and the label selection area includes the plurality of label controls. As Figure 3 shown, step S202 may include: Step S300, receiving the text content through the text input area, where the text content includes a plurality of text paragraphs.

[0051] Step S302, in response to selecting any one of the text paragraphs and a target label type selected from the plurality of label types for the text paragraph, creating a label corresponding to the target label type after the text paragraph; or in response to selecting any two text paragraphs and a target label type selected from the plurality of label types for the two text paragraphs, creating a label corresponding to the target label type between the two text paragraphs.

[0052] Step S304, in response to selecting the label, presenting a label attribute adjustment window, where the label attribute adjustment window includes a plurality of label attributes or label attribute input boxes corresponding to the target label type.

[0053] Step S306, in response to selecting one of the label attributes, determining the selected label attribute as the target label attribute, or receiving the target label attribute through the label attribute input box.

[0054] Combined with Figure 2 , the text content can be received in the text input area by inputting text or pasting data, etc. In some embodiments, when receiving the text content by pasting data, the pasted data can be cleaned. Specifically, the system-level clipboard reading function can be called through "navigator.clipboard.read()" to asynchronously obtain the data in the clipboard. For example, a MIME (Multipurpose Internet Mail Extensions) type priority queue can be established to preferentially parse data in the text / html format and degrade and be compatible with data in the text / plain format to improve the integrity and availability of the data. During the parsing process, DOMParser is used to implement non-rendering environment parsing, and a DOM structure purifier is pre-constructed to filter out risk labels in the data based on the whitelist mechanism (such as <script>、<iframe>、onclick等),防止XSS攻击,确保粘贴数据的安全性。

[0055] 在实际应用中,可以针对不同的文本段落定制不同的标签类型和标签属性。比如,若用户选择"hello”,再从标签选择区域中选择插入停顿控件。在所选"hello”的段落后创建与目标标签类型对应的"停顿”标签。用户点击"停顿”标签,编辑界面可以弹出标签属性调整窗口,以供用户设置停顿时间。若用户选定多音字选择控件时,首先通过isHanzi方法检测当前选中文本段落中的汉字字符,具体包括:遍历文本的Unicode编码并查询预加载的拼音字典(pinyin_dict)的键集合,标记出所有存在于拼音字典中的汉字字符。随后,识别出这些多音字字符的每个多音字字符。调用getPinyin方法从拼音字典中获取原始拼音列表(如"中”字符对应["zhōng”,"zhòng”]),并通过声调映射矩阵(toneMap)将其转换为数字声调形式(如"zhōng”转为"zhong1”)。在一些实施例中,还可以自动扩展无调号基础形式(如补充"zhong”)。在一些实施例中,还可以通过键存在性检测机制替代正则表达式匹配,能够显著提升汉字字符的鉴别效率。最终可以确定包含原始拼音、数字声调形式及其无调号变体的完整拼音结果集,并将拼音结果集可视化呈现在编辑界面的标签属性调整窗口中。需要说明的是,本申请实施例通过"原始拼音+数字声调+无调号形式”的组合输出策略,兼容不同SSML引擎的拼音输入规范。通过前端预置拼音字典,可以降低网络请求依赖,提高实时多音字处理的服务稳定性。通过哈希表可以实现O(1)复杂度的声调分解。

[0056] 在一些实施例中,在用户输入的标签属性不符合规范的情况下,可以在编辑界面显示错误提示。比如,用户输入time="500”(缺少单位),在编辑界面提示"请输入正确的停顿时长(如500ms)。

[0057] 在本实施例中,用户可在文本输入区域输入多段文本,并通过标签选择区域为单个段落添加指定类型的标签,或在段落之间设置指定类型的标签,同时支持通过弹窗动态调整标签属性(即选择预设好的标签属性或输入自定义标签属性),从而提升文本标注的灵活性和效率。

[0058] 步骤S204,根据所述文本内容、所述目标标签类型和目标标签属性,在所述DOM结构中创建文本节点和标签节点。

[0059] 当用户输入或粘贴文本内容时,可以通过document.createTextNode()生成文本节点。将生成的文本节点插入到DOM树的对应位置,以实时反映在编辑界面中。

[0060] 根据用户选定的目标标签类型(如<break>、<phoneme>),可以调用document.createElement()生成对应标签节点,并通过setAttribute()注入配置的标签属性(如time="500ms”)。在一些实施例中,还可以校验标签属性的合法性(如单位、格式)。

[0061] 在可选的实施例,如图4所示,步骤S204可以包括:步骤S400,根据所述文本段落,在所述DOM结构中创建对应的文本节点。

[0062] 步骤S402,根据每个标签对应的目标标签类型和目标标签属性,在所述DOM结构中创建对应的标签节点。其中,根据所述文本段落与所述标签的位置关系,将所述标签节点添加到对应的文本节点后或者插入对应的两个文本节点之间。

[0063] 示例性地,可以对输入的文本内容进行解析,根据自然语言特征或用户配置的标注粒度,自动将文本划分为具有语义完整性的文本段落。

[0064] 所述文本节点包括原始文本内容、位置信息和上下文关系等结构化元数据。所述标签节点包括标签类型标识(如<phoneme> / <break>的DOM元素类型)、配置的标签属性(如alphabet="pinyin”、time="200ms”等)、作用范围标记(记录其关联的文本节点起止位置)、渲染控制参数(如可视化样式、交互状态等UI相关属性)等。

[0065] 在本实施例中,通过解析文本段落与标签的关联关系,在DOM结构中动态创建文本节点和标签节点,并基于位置关系(如相邻或嵌套)将标签节点精准插入到文本节点的指定位置,减少手动操作调整DOM结构的繁琐步骤。基于预设规则自动生成节点,降低开发成本。

[0066] 步骤S206,将所述DOM结构中的文本节点转换为SSML文本,以及将所述DOM结构中的标签节点转换为SSML标签。

[0067] 示例性地,可以根据实际需求,从文本节点中提取有效信息转换成SSML文本,过滤掉无效字符,提升SSML的整体质量。例如:可以通过getTextWithLineBreaks递归提取文本节点中的纯文本内容。将纯文本内容作为SSML文本。同样地,可以根据实际需求,将DOM结构中地标签节点一一转换成类型对应的SSML标签。

[0068] 可选的实施例中,如图5所示,步骤S206可以包括:步骤S500,创建SSML容器。

[0069] 步骤S502,根据所述文本节点和所述标签节点在所述DOM结构中的排序,在所述SSML容器中依次生成对应的SSML文本和SSML标签。其中,对于每个文本节点,通过正则表达式从对应的文本段落中提取有效文本,并将所述有效文本转换为所述SSML文本;对于每个标签节点,从所述标签节点中提取出目标标签类型和目标标签属性,并将所述目标标签类型和目标标签属性转换为所述SSML标签。

[0070] 在本实施例中,根据各个节点(文本节点和标签节点)在DOM结构中的顺序,在SSML容器中依次生成SSML文本和SSML标签,节省人工编写代码的成本。

[0071] 步骤S208,根据所述SSML文本和所述SSML标签合成SSML数据,所述SSML数据用于指导针对所述文本内容的语音合成。

[0072] 示例性地,通过XMLSerializer序列化SSML文本和SSML标签,生成标准的SSML字符串(SSML数据)。将SSML数据发送到TTS(文本转语音系统),可以将文本内容转换为语音,语音效果与SSML标签相匹配。

[0073] 在一些实施例中,还可以通过规则引擎对SSML的语法进行校验,例如检查标签嵌套的合法性(如禁止<phoneme>嵌套<say-as>)以及检测必填属性的缺失情况(如<prosody>缺少rate属性),从而提高合成SSML数据的规范性。

[0074] 在一些实施例中,在预览SSML合成语音效果时,可以采用分层式审核机制进行双重合规检测。具体地,首先执行文本预审,通过解析SSML结构提取待合成文本进行多模态分析。随后可以对通过文本审核的内容生成低码率预览音频,执行声学合规复检。该机制通过独立构建的内容安全中间件实现,其核心模块包括:(1)语义分析引擎,基于BERT+BiLSTM混合网络识别文本违规内容,并集成领域知识图谱实现动态词库扩展;(2)声学特征检测器,通过提取音频MFCC特征并应用CNN网络检测不合规声学信号,同时利用声纹比对技术阻断仿冒声线请求;(3)策略管理中心,负责多维度审核规则的动态配置与实时更新。

[0075] 在可选的实施例中,如图6所示,所述SSML编辑方法还可以包括:步骤S600,根据所述SSML数据,生成对应的历史编辑记录;步骤S602,在所述编辑界面展示所述历史编辑记录;步骤S604,响应于选定所述历史编辑记录,在所述编辑界面展示所述文本内容以及针对所述文本内容选定的目标标签类型和目标标签属性。

[0076] 所述历史编辑记录可以包括完整的SSML数据、时间戳、操作前后的DOM结构快照、用户标识等。可以通过时间范围查询、操作类型查询、关键词搜索、标签类型检索等各种方式查询历史编辑记录。在查询到符合条件的历史编辑记录后,可以通过时间轴视图、缩略图预览、标签云等各种方式展示历史编辑记录。可以通过悬停查看、点击应用按钮(比如,一键还原按钮)、拖拽等各种方式选定历史编辑记录。

[0077] 在本实施例中,通过记录SSML编辑历史并可视化展示,实现编辑过程的可追溯与快速回退,提升编辑效率和容错能力。

[0078] 在可选的实施例中,如图7所示,步骤S604可以包括:步骤S700,响应于选定所述历史编辑记录,获取对应的SSML数据。

[0079] 步骤S702,通过预先构建好的DOM解析器对SSML数据进行解析,并转换为DOM结构。其中,所述SSML数据包括SSML文本和SSML标签,所述DOM结构包括文本节点和标签节点,所述SSML文本对应所述文本节点,所述SSML标签对应所述标签节点。

[0080] 步骤S704,根据所述DOM结构,在所述编辑界面渲染出所述文本内容以及针对所述文本内容选定的目标标签类型和目标标签属性。

[0081] 示例性地,可以通过预构建的DOM解析器对历史SSML数据进行语法解析,提取SSML文本、SSML标签、以及SSML文本和SSML标签的嵌套关系等数据。将解析得到的数据转换为DOM节点时,可以为每个SSML标签创建带语义标注的SPAN容器元素。通过JSON序列化将SSML标签属性对应的元数据存储在HTML元素的data-props特性中。在SPAN元素的dataset属性中完整保存原始SSML标签名、SSML标签属性键值对、关联的SSML文本等。在一些实施例中,可以通过CSS样式元素和背景色块来区分不同的标签。在一些实施例中,为每个可视化标签集成右键菜单系统。在一些实施例中,可以建立DOM结构与文本输入框中内容(文本内容和标签)的双向映射关系,使用MutationObserver监控DOM变化,实时更新对应的文本内容和标签,从而实现版本对比功能。

[0082] 在本实施例中,通过解析历史SSML数据并重建DOM结构,实现编辑内容的精准还原与可视化呈现,显著提升版本回溯效率和编辑准确性。

[0083] 在可选的实施例中,步骤S702可以包括:从所述SSML标签中提取目标标签类型和目标标签属性。获取所述目标标签类型对应的多个标签属性,根据所述目标标签类型、目标标签属性和所述多个标签属性,创建所述标签节点。相应地,步骤S704包括:根据所述文本节点和所述标签节点,在所述编辑界面渲染多个文本段落以及对应的标签。其中,每个标签携带所述目标标签属性。在所述标签被选定的情况下,所述编辑界面展示标签属性调整窗口,所述标签属性调整窗口包括所述多个标签属性。

[0084] 示例性地,可以通过扫描DOM树定位带有ssml-label类名的SPAN元素,解析其data-props属性中的JSON格式元数据,以提取目标标签类型和目标标签属性等。然后,基于预设的可扩展SSML标签映射规则,获取目标标签类型对应的多个标签属性。然后,根据目标标签类型、目标标签属性和多个标签属性,创建标签节点。

[0085] 在本实施例中,通过解析SSML标签提取目标标签类型及目标标签属性,将其转换为携带完整语义信息的标签节点(HTML标签节点),同时结合文本节点(HTML文本节点)在编辑界面进行可视化渲染,实现了从结构化SSML数据到HTML可视化元素的逆向转换。

[0086] 为了使得本申请更加容易理解,以下结合图8A、图8B、图9和图10提供一个示例性应用。

[0087] 本申请实施例用于SSML编辑器,所述SSML编辑器包括编辑器主控件、标签控件、顶部操作按钮、面板控件(文本替换面板控件、下拉选择面板控件)、试听按钮等。

[0088] 步骤S11,显示基于DOM结构的编辑界面。其中,所述编辑页面包括文本输入区域和标签选择区域。所述标签选择区域包括多个标签控件。多个标签控件包括多音字控件、数字读法控件、发音替换控件等。一个标签控件对应一个标签类型。

[0089] 步骤S12,通过文本输入区域接收文本内容,所述文本内容包括多个文本段落。若所述文本内容长度超过预设长度,删除部分文本。

[0090] 步骤S13,在选择文本内容中的部分文本段落,并点击标签选择区域中的目标标签控件(即选定目标标签类型),为选中的文本段落创建对应的标签。

[0091] 具体地,响应于选定任一个文本段落,在文本段落后创建与目标标签类型对应的标签;或者响应于选中任意两个文本段落,在所述两个文本段落之间创建与目标标签类型对应的标签。

[0092] 步骤S14,响应于点击标签,弹出标签对应的标签属性调整窗口以接收目标标签属性。

[0093] 具体地,所述标签属性调整窗口包括所述目标标签类型对应的多个标签属性或标签属性输入框。响应于选定其中一个标签属性,将所选定的标签属性确定为目标标签属性,或者通过标签属性输入框接收目标标签属性。

[0094] 步骤S15,根据文本内容、目标标签类型和目标标签属性,在DOM结构中创建文本节点和标签节点。

[0095] 步骤S16,将所述DOM结构中的文本节点转换为SSML文本,以及将所述DOM结构中的标签节点转换为SSML标签。

[0096] 步骤S17,根据SSML文本和SSML标签合成SSML数据。

[0097] 步骤S18,响应于触发试听按钮,将所述SSML数据传递到传递到TTS服务中,生成对应针对文本内容的语音(比如MP3文件)。

[0098] 步骤S19,响应于触发保存按钮,根据所述SSML数据,生成对应的历史编辑记录。

[0099] 步骤S20,在编辑界面展示历史编辑记录,并响应于选定历史编辑记录,在编辑界面展示文本内容以及针对文本内容的标签。其中,在标签被选定的情况下,在编辑界面中展示标签属性调整窗口,所述标签属性调整窗口包括标签对应的多个标签属性。

[0100] 在该示例性应用中,通过可视化界面简化SSML标签的编辑,使用户能够更轻松地插入、删除和调整标签,无需手动编写复杂的代码,极大地提高了编辑效率。在编辑过程中,可以实时听到语音合成效果。通过将标签与Vue框架的属性和事件绑定,实现SSML标记的动态交互。提供字数限制功能和多音字识别,以尽量确保生成的SSML数据符合语音合成引擎的要求,还可以实现SSML格式和HTML格式的双向转换,使得数据在编辑、导入和导出过程中的格式兼容性问题得到有效解决。

[0101] 实施例二图11示意性示出了根据本申请实施例二的SSML编辑装置的框图,该装置可以被分割成一个或多个程序模块,一个或者多个程序模块被存储于存储介质中,并由一个或多个处理器所执行,以完成本申请实施例。本申请实施例所称的程序模块是指能够完成特定功能的一系列计算机程序指令段,以下描述将具体介绍本实施例中各程序模块的功能。如图11所示,该装置1000可以包括:显示模块1100、接收模块1200、创建模块1300、转换模块1400、合成模块1500,其中:显示模块1100,用于显示基于DOM结构的编辑界面,所述编辑界面配置有多个标签控件,多个控件对应多个标签类型,每个标签类型具有多个标签属性;接收模块1200,用于通过所述编辑界面接收文本内容以及确定针对所述文本内容选定的目标标签类型和目标标签属性;创建模块1300,用于根据所述文本内容、所述目标标签类型和目标标签属性,在所述DOM结构中创建文本节点和标签节点;转换模块1400,用于将所述DOM结构中的文本节点转换为SSML文本,以及将所述DOM结构中的标签节点转换为SSML标签;合成模块1500,用于根据所述SSML文本和所述SSML标签合成SSML数据,所述SSML数据用于指导针对所述文本内容的语音合成。

[0102] 在可选的实施例中,所述编辑界面包括文本输入区域和标签选择区域,所述标签选择区域包括所述多个标签控件;所述接收模块1200用于:通过所述文本输入区域接收所述文本内容,所述文本内容包括多个文本段落;响应于选定任一个所述文本段落以及为所述文本段落从所述多个标签类型中选择的目标标签类型,在所述文本段落后创建与所述目标标签类型对应的标签;或者响应于选中任意两个文本段落以及为所述两个文本段落从所述多个标签类型中选择的目标标签类型,在所述两个文本段落之间创建与所述目标标签类型对应的标签;响应于选定所述标签,展示标签属性调整窗口,所述标签属性调整窗口包括所述目标标签类型对应的多个标签属性或标签属性输入框;响应于选定其中一个标签属性,将所选定的标签属性确定为所述目标标签属性,或者通过所述标签属性输入框接收所述目标标签属性。

[0103] 在可选的实施例中,创建模块1300还用于:根据所述文本段落,在所述DOM结构中创建对应的文本节点;根据每个标签对应的目标标签类型和目标标签属性,在所述DOM结构中创建对应的标签节点;其中,根据所述文本段落与所述标签的位置关系,将所述标签节点添加到对应的文本节点后或者插入对应的两个文本节点之间。

[0104] 在可选的实施例中,转换模块1400还用于:创建SSML容器;根据所述文本节点和所述标签节点在所述DOM结构中的排序,在所述SSML容器中依次生成对应的SSML文本和SSML标签;其中,对于每个文本节点,通过正则表达式从对应的文本段落中提取有效文本,并将所述有效文本转换为所述SSML文本;对于每个标签节点,从所述标签节点中提取出目标标签类型和目标标签属性,并将所述目标标签类型和目标标签属性转换为所述SSML标签。

[0105] 在可选的实施例中,所述SSML编辑装置还包括展示模块,用于:根据所述SSML数据,生成对应的历史编辑记录;在所述编辑界面展示所述历史编辑记录;响应于选定所述历史编辑记录,在所述编辑界面展示所述文本内容以及针对所述文本内容选定的目标标签类型和目标标签属性。

[0106] 在可选的实施例中,响应于选定所述历史编辑记录,在所述编辑界面展示所述文本内容以及针对所述文本内容选定的目标标签类型和目标标签属性,包括:响应于选定所述历史编辑记录,获取对应的SSML数据;通过预先构建好的DOM解析器对SSML数据进行解析,并转换为DOM结构;其中,所述SSML数据包括SSML文本和SSML标签,所述DOM结构包括文本节点和标签节点,所述SSML文本对应所述文本节点,所述SSML标签对应所述标签节点;根据所述DOM结构,在所述编辑界面渲染出所述文本内容以及针对所述文本内容选定的目标标签类型和目标标签属性。

[0107] 在可选的实施例中,通过预先构建好的DOM解析器对SSML数据进行解析,并转换为DOM结构,包括:从所述SSML标签中提取目标标签类型和目标标签属性;获取所述目标标签类型对应的多个标签属性,根据所述目标标签类型、目标标签属性和所述多个标签属性,创建所述标签节点;对应地,根据所述DOM结构,在所述编辑界面渲染出所述文本内容以及针对所述文本内容选定的目标标签类型和目标标签属性,包括:根据所述文本节点和所述标签节点,在所述编辑界面渲染多个文本段落以及对应的标签;其中,每个标签携带所述目标标签属性;在所述标签被选定的情况下,所述编辑界面展示标签属性调整窗口,所述标签属性调整窗口包括所述多个标签属性。

[0108] 实施例三图12示意性示出了根据本申请实施例三的适于实现SSML编辑方法的计算机设备10000的硬件架构示意图。在一些实施例中,计算机设备10000可以是智能手机、可穿戴设备、平板电脑、个人电脑、车载终端、游戏机、虚拟设备、工作台、数字助理、机顶盒、机器人等终端设备。在另一些实施例中,计算机设备10000可以是机架式服务器、刀片式服务器、塔式服务器或机柜式服务器(包括独立的服务器,或多个服务器所组成的服务器集群)等。如图12所示,所述计算机设备10000包括但不限于:可通过系统总线相互通信链接存储器10010、处理器10020、网络接口10030。其中:存储器10010至少包括一种类型的计算机可读存储介质,可读存储介质包括闪存、硬盘、多媒体卡、卡型存储器(如,SD或DX存储器)、随机访问存储器(RAM)、静态随机访问存储器(SRAM)、只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、可编程只读存储器(PROM)、磁性存储器、磁盘、光盘等。在一些实施例中,存储器10010可以是计算机设备10000的内部存储模块,例如该计算机设备10000的硬盘或内存。在另一些实施例中,存储器10010也可以是计算机设备10000的外部存储设备,例如该计算机设备10000上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)等。当然,存储器10010还可以既包括计算机设备10000的内部存储模块也包括其外部存储设备。本实施例中,存储器10010通常用于存储安装于计算机设备10000的操作系统和各类应用软件,例如SSML编辑方法的程序代码等。此外,存储器10010还可以用于暂时地存储已经输出或者将要输出的各类数据。

[0109] 处理器10020在一些实施例中可以是中央处理器(Central Processing Unit,CPU)、控制器、微控制器、微处理器、或其他芯片。该处理器10020通常用于控制计算机设备10000的总体操作,例如执行与计算机设备10000进行数据交互或者通信相关的控制和处理等。本实施例中,处理器10020用于运行存储器10010中存储的程序代码或者处理数据。

[0110] 网络接口10030可包括无线网络接口或有线网络接口,该网络接口10030通常用于在计算机设备10000与其他计算机设备之间建立通信链接。例如,网络接口10030用于通过网络将计算机设备10000与外部终端相连,在计算机设备10000与外部终端之间建立数据传输通道和通信链接等。网络可以是企业内部网(Intranet)、互联网(Internet)、全球移动通讯系统(Global System of Mobile communication,简称为GSM)、宽带码分多址(WidebandCode Division Multiple Access,简称为WCDMA)、4G网络、5G网络、蓝牙(Bluetooth)、Wi-Fi等无线或有线网络。

[0111] 需要指出的是,图12仅示出了具有部件10010-10030的计算机设备,但是应该理解的是,并不要求实施所有示出的部件,可以替代地实施更多或者更少的部件。

[0112] 在本实施例中,存储于存储器10010中的SSML编辑方法还可以被分割为一个或者多个程序模块,并由一个或多个处理器(如处理器10020)所执行,以完成本申请实施例。

[0113] 实施例四本申请实施例还提供一种计算机可读存储介质,计算机可读存储介质其上存储有计算机程序,其中,计算机程序被处理器执行时实现实施例中的SSML编辑方法的步骤。

[0114] 本实施例中,计算机可读存储介质包括闪存、硬盘、多媒体卡、卡型存储器(例如,SD或DX存储器等)、随机访问存储器(RAM)、静态随机访问存储器(SRAM)、只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、可编程只读存储器(PROM)、磁性存储器、磁盘、光盘等。在一些实施例中,计算机可读存储介质可以是计算机设备的内部存储单元,例如该计算机设备的硬盘或内存。在另一些实施例中,计算机可读存储介质也可以是计算机设备的外部存储设备,例如该计算机设备上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)等。当然,计算机可读存储介质还可以既包括计算机设备的内部存储单元也包括其外部存储设备。本实施例中,计算机可读存储介质通常用于存储安装于计算机设备的操作系统和各类应用软件,例如实施例中SSML编辑方法的程序代码等。此外,计算机可读存储介质还可以用于暂时地存储已经输出或者将要输出的各类数据。

[0115] 实施例五本申请实施例还提供一种计算机程序产品,包括计算机程序,该计算机程序被处理器执行时实现上述实施例中的方法。

[0116] 显然,本领域的技术人员应该明白,上述的本申请实施例的各模块或各步骤可以用通用的计算机设备来实现,它们可以集中在单个的计算机设备上,或者分布在多个计算机设备所组成的网络上,可选地,它们可以用计算机设备可执行的程序代码来实现,从而,可以将它们存储在存储装置中由计算机设备来执行,并且在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤,或者将它们分别制作成各个集成电路模块,或者将它们中的多个模块或步骤制作成单个集成电路模块来实现。这样,本申请实施例不限制于任何特定的硬件和软件结合。

[0117] 需要说明的是,以上仅为本申请的优选实施例,并非因此限制本申请的专利保护范围,凡是利用本申请说明书及附图内容所作的等效结构或等效流程变换,或直接或间接运用在其他相关的技术领域,均同理包括在本申请的专利保护范围内。< / script> < / phoneme>

Claims

1. An SSML editing method, characterized in that, The method includes: Displaying an editing interface based on a DOM structure, the editing interface being configured with a plurality of tag controls, the plurality of controls corresponding to a plurality of tag types, and each tag type having a plurality of tag attributes; Receiving text content through the editing interface and determining a target tag type and target tag attributes selected for the text content; Creating a text node and a tag node in the DOM structure according to the text content, the target tag type, and the target tag attributes; Converting the text node in the DOM structure into SSML text, and converting the tag node in the DOM structure into an SSML tag; Synthesizing SSML data according to the SSML text and the SSML tag, the SSML data being used to guide speech synthesis for the text content.

2. The method according to claim 1, characterized in that, The editing interface includes a text input area and a tag selection area, the tag selection area including the plurality of tag controls; receiving text content through the editing interface and determining a target tag type and target tag attributes selected for the text content includes: Receiving the text content through the text input area, the text content including a plurality of text paragraphs; In response to selecting any one of the text paragraphs and a target tag type selected from the plurality of tag types for the text paragraph, creating a tag corresponding to the target tag type after the text paragraph; or in response to selecting any two text paragraphs and a target tag type selected from the plurality of tag types for the two text paragraphs, creating a tag corresponding to the target tag type between the two text paragraphs; In response to selecting the tag, presenting a tag attribute adjustment window, the tag attribute adjustment window including a plurality of tag attributes or tag attribute input boxes corresponding to the target tag type; In response to selecting one of the tag attributes, determining the selected tag attribute as the target tag attribute, or receiving the target tag attribute through the tag attribute input box.

3. The method according to claim 2, wherein Creating a text node and a tag node in the DOM structure according to the text content, the target tag type, and the target tag attributes includes: Creating a corresponding text node in the DOM structure according to the text paragraph; Creating a corresponding tag node in the DOM structure according to the target tag type and target tag attributes corresponding to each tag; Wherein, according to the positional relationship between the text paragraph and the tag, the tag node is added after the corresponding text node or inserted between the corresponding two text nodes.

4. The method according to claim 3, wherein Converting the text node in the DOM structure into SSML text, and converting the tag node in the DOM structure into an SSML tag includes: Creating an SSML container; Generating corresponding SSML text and SSML tags in sequence in the SSML container according to the sorting of the text node and the tag node in the DOM structure; Among them, for each text node, valid text is extracted from the corresponding text paragraph through regular expressions, and the valid text is converted into the SSML text; for each tag node, the target tag type and target tag attributes are extracted from the tag node, and the target tag type and target tag attributes are converted into the SSML tag.

5. The method according to claim 2, characterized in that, The method further includes: generating a corresponding historical edit record according to the SSML data; displaying the historical edit record on the edit interface; in response to selecting the historical edit record, displaying the text content and the selected target tag type and target tag attributes for the text content on the edit interface.

6. The method according to claim 5, characterized in that, In response to selecting the historical edit record, displaying the text content and the selected target tag type and target tag attributes for the text content on the edit interface includes: in response to selecting the historical edit record, obtaining the corresponding SSML data; parsing the SSML data through a pre-constructed DOM parser and converting it into a DOM structure; wherein, the SSML data includes SSML text and SSML tags, the DOM structure includes text nodes and tag nodes, the SSML text corresponds to the text nodes, and the SSML tags correspond to the tag nodes; rendering the text content and the selected target tag type and target tag attributes for the text content on the edit interface according to the DOM structure.

7. The method according to claim 6, characterized in that, Parsing the SSML data through a pre-constructed DOM parser and converting it into a DOM structure includes: extracting the target tag type and target tag attributes from the SSML tags; obtaining multiple tag attributes corresponding to the target tag type, and creating the tag node according to the target tag type, target tag attributes, and the multiple tag attributes; Correspondingly, rendering the text content and the selected target tag type and target tag attributes for the text content on the edit interface according to the DOM structure includes: rendering multiple text paragraphs and corresponding tags on the edit interface according to the text nodes and the tag nodes; wherein each tag carries the target tag attributes; when the tag is selected, a tag attribute adjustment window is displayed on the edit interface, and the tag attribute adjustment window includes the multiple tag attributes.

8. An SSML editing device, characterized in that, The device includes: a display module for displaying an edit interface based on the DOM structure, the edit interface being configured with multiple tag controls, the multiple controls corresponding to multiple tag types, and each tag type having multiple tag attributes; a receiving module for receiving the text content and determining the selected target tag type and target tag attributes for the text content through the edit interface; a creating module for creating text nodes and tag nodes in the DOM structure according to the text content, the target tag type, and the target tag attributes; a conversion module for converting the text nodes in the DOM structure into SSML text and converting the tag nodes in the DOM structure into SSML tags; A synthesis module, configured to synthesize SSML data according to the SSML text and the SSML tags, where the SSML data is used to guide speech synthesis for the text content.

9. A computer device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein: The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Computer instructions are stored in the computer-readable storage medium, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to claims 1 to 7 are implemented.

Citation Information

Cited By

  • Voice generation method and device, storage medium and electronic equipment

    CN120808747A