Text processing method and device, equipment, medium and product
By recognizing and renaming the audio files, the problem of inaccurate matching between audio and dialogue was solved, achieving efficient and accurate synchronous output of audio and dialogue, thus improving the user experience.
Patent Information
- Application Number
- CN202511404524.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2026-01-16
AI Technical Summary
In existing technologies, the matching of voice resources and game lines relies on manual adjustments, which is costly and prone to errors, resulting in mismatches between voice and lines.
By acquiring the audio file, performing speech recognition processing to generate audio text, and matching the target dialogue text with the pre-acquired dialogue text, the audio file is renamed based on the text identifier of the target dialogue text to ensure that the audio file and the dialogue text have the same identifier.
It improves the efficiency and accuracy of voice and dialogue matching, reduces costs, ensures the accuracy of synchronous output of voice and dialogue, and enhances the readability and reliability of information obtained by users.
Smart Images

Figure CN121350221A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer processing technology, and in particular to a text processing method, apparatus, device, medium, and product. Background Technology
[0002] In the development of digital games, it is often necessary to add corresponding voice resources for various characters in the game's storyline. To enable users to better understand the voice content, it is usually necessary to configure corresponding game lines for the voice resources so that the corresponding game lines are displayed when the voice resources are played.
[0003] In existing technologies, the processing of voice resources and game lines usually relies on manual comparison of the content in the game lines and voice resources, and manual correction of the file identifiers of the voice or lines so that the voice file identifier and the game lines with the same name are the same, so that when the voice resources are played, the game lines with the same name as the voice are displayed.
[0004] This method, which relies on manual adjustment of names, requires manual comparison of each item, which is costly and prone to errors, leading to mismatches between voice and dialogue. Summary of the Invention
[0005] This invention provides a text processing method, apparatus, device, medium, and product to improve the efficiency and accuracy of matching voice files with dialogue text, and to ensure the matching accuracy between voice and dialogue.
[0006] According to one aspect of the present invention, a file processing method is provided, the method comprising:
[0007] Obtain at least one audio file;
[0008] For each of the aforementioned audio files, the audio file is processed to obtain the audio text corresponding to the audio file;
[0009] From at least one pre-acquired dialogue text to be dubbed, determine the target dialogue text that matches the audio text;
[0010] Based on the text identifier of the target dialogue text, the file identifier of the audio file is renamed to synchronously output the audio file and the target dialogue text with the same identifier.
[0011] According to another aspect of the present invention, a document processing apparatus is provided, the apparatus comprising:
[0012] The audio file acquisition module is used to acquire at least one audio file;
[0013] The speech-to-text determination module is used to process each speech file to obtain the speech-to-text corresponding to the speech file.
[0014] The target dialogue text determination module is used to determine the target dialogue text that matches the voice text from at least one pre-acquired dialogue text to be dubbed;
[0015] The renaming module is used to rename the file identifier of the audio file based on the text identifier of the target dialogue text, so as to synchronously output the audio file and the target dialogue text with the same identifier.
[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0017] At least one processor; and a memory communicatively connected to said at least one processor; wherein,
[0018] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the file processing method according to any embodiment of the present invention.
[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the file processing method described in any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the file processing method as described in any embodiment of the present invention.
[0021] The technical solution of this invention involves acquiring at least one audio file; processing each audio file to obtain corresponding audio text; determining a target dialogue text matching the audio text from at least one pre-acquired dialogue text to be dubbed; and renaming the file identifier of the audio file based on the text identifier of the target dialogue text to synchronously output audio files and target dialogue texts with the same identifier. This solves the problem of high cost and easy matching errors in the existing technology, which relies on manual renaming, leading to mismatches between audio and dialogue. The solution achieves the acquisition of audio text corresponding to each audio file through recognition. Matching the acquired at least one dialogue text to be dubbed and the audio text to determine the target dialogue text improves the efficiency and accuracy of audio-text matching. Furthermore, renaming the file identifier of the audio file based on the text identifier of the target dialogue text enhances the efficiency and accuracy of matching audio files with dialogue texts, reduces costs, and ensures the matching accuracy between synchronously output audio files with the same identifier and target dialogue texts, thereby improving the readability and reliability of information obtained by users.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of a file processing method provided according to Embodiment 1 of the present invention;
[0025] Figure 2 This is a flowchart of a file processing method provided according to Embodiment 2 of the present invention;
[0026] Figure 3 This is a schematic diagram of the structure of a document processing device according to Embodiment 3 of the present invention;
[0027] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the file processing method of the present invention. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solution disclosed herein all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to maintain user personal information security and network security. It should also be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solution disclosed herein are all conducted with the user's knowledge and consent, and comply with relevant privacy protection regulations.
[0031] Example 1
[0032] Figure 1 This is a flowchart of a file processing method according to Embodiment 1 of the present invention. This embodiment is applicable to any situation requiring the synchronous output of dialogue text adapted to an audio file. This method can be executed by a file processing device, which can be implemented in hardware and / or software and can be configured in a computing device. Figure 1 As shown, the method includes:
[0033] S110. Obtain at least one audio file.
[0034] Voice files refer to audio data stored digitally, used to record user voice input or system-generated voice output. Voice files can be audio files in various formats such as WAV, MP3, AAC, OGG, and PCM. For example, in a game scenario, voice files can include voiceovers for game characters, voice commands (such as "forward" or "gather"), voices of non-player characters (NPCs) (such as background stories or instructions for the next action), audio of various sound effects (such as wind, rain, and birdsong), game prompts, background music, or dynamic voice content generated through machine learning models.
[0035] In embodiments of the present invention, methods for acquiring at least one audio file include, but are not limited to: responding to a user's voice input operation triggered by a voice input device (such as a microphone, a voice interface of a mobile terminal, etc.), real-time acquisition of the user's dictated voice, and generating an audio file; reading a stored audio file from a local storage device; downloading an audio file from a remote server via a network interface; automatically generating an audio file based on a game system according to preset rules; and converting text information into an audio file using a speech synthesizer. It should be noted that the total number of acquired audio files can be one or more, and each audio file can represent a segment of audio content.
[0036] In this embodiment of the invention, multiple audio files can be audio files determined by different data sources. The file identifier naming methods differ between different data sources. That is, the at least one audio file may include audio files determined by different data sources. Different data sources use different file identifier naming methods to mark the file identifiers of the audio files they determine.
[0037] The data source can refer to the source providing the audio files. Optionally, the data source may include at least a voice input device (such as a recording device, online voice service, etc.), a voice generation model, and a voice synthesizer. The file identification naming method refers to the naming rules used to uniquely identify and distinguish different audio files. For example, the file identification naming method for voice input devices can be based on a combination of information elements such as timestamps, source type, and user ID. The file identification naming method for voice generation models can be based on a combination of information elements such as model version number, generation time, and semantic tags. The file identification naming method for voice synthesizers can be based on a combination of information elements such as text content summary, synthesis time, and timbre selection.
[0038] Furthermore, the file identification naming method is not limited to the information elements mentioned above, and can also be combined with other relevant information elements, such as device model, network status, etc.
[0039] Optionally, the audio file is determined based on at least one of the following methods:
[0040] One implementation method is to acquire at least one first speech segment based on a voice input device, obtain a speech file, and name the speech file.
[0041] The first audio segment can be a user's voice fragment collected in real time through a voice input device. For example, the first audio segment can be the voice content spoken by the user or the dialogue content of multiple users.
[0042] In this embodiment of the invention, user-inputted speech can be collected in real time using a voice input device to form one or more first speech segments. Each first speech segment is processed and saved as a speech file. For ease of management and subsequent use, the voice input device can assign a unique file identifier to each speech file according to a built-in file identifier naming method. For example, the file identifier may include information elements such as user ID, collection time, and speech length.
[0043] For example, the voice input device is a recording device. When a user speaks, the recording control in the voice input device can be triggered to start recording. When recording is complete, the recording end control can be triggered to encapsulate the recorded voice into a voice file and name the file identifier of the voice file. It should be noted that the file identifier naming method may be different for different voice input devices, which will not be elaborated here.
[0044] Another approach is to generate at least one speech file based on a speech generation model and name the speech file.
[0045] The speech generation model can be a pre-trained model used to generate speech files with specific semantics.
[0046] In this embodiment of the invention, textual or semantic information of the speech to be generated can be input into the speech generation model. The speech generation model outputs corresponding speech content based on the given semantic or textual information. The speech generation model can encapsulate the speech content into a speech file and name the speech file's file identifier based on the file identifier naming method in the speech generation model. For example, the file identifier can include information elements such as model version number, generation time, and semantic tags.
[0047] Another approach is to synthesize at least one second speech segment using a speech synthesizer to obtain a speech file, and then name the speech file.
[0048] Among them, a speech synthesizer can be software, application, tool or model that combines multiple speech segments into a single speech file.
[0049] In this embodiment of the invention, at least one second speech segment to be synthesized can be input into a speech synthesizer to output a speech file. The speech synthesizer names the output speech file with a file identifier. For example, the file identifier may include information such as text summary, synthesis time, and timbre selection.
[0050] By using the above methods to determine the speech file, we can improve the flexibility, efficiency, and diversity of speech file determination, and meet the needs of generating diverse speech.
[0051] In practical applications, since different data sources use different file identification naming methods for audio files, directly mapping the file identifiers of audio files to the identifiers of the dialogue text in the script list results in complex operations, poor matching accuracy, and low efficiency. To solve this problem, file processing can be performed based on the technical solution provided in this embodiment of the invention, mapping the identifiers of audio files and dialogue text one-to-one, thereby improving processing efficiency and matching accuracy.
[0052] S120. For each audio file, process the audio file to obtain the audio text corresponding to the audio file.
[0053] Among them, speech text can be the textual representation of the speech content in a speech file, which reflects the semantic information contained in the speech file.
[0054] In this embodiment of the invention, each audio file can be processed by speech recognition to extract its semantic content, and then audio text corresponding to that audio file can be generated based on the semantic content. Alternatively, for each acquired audio file, the audio file can be input into a semantic recognition module for processing, such as audio feature extraction, speech endpoint detection, acoustic model matching, and language model parsing, to generate audio text corresponding to the audio file.
[0055] For example, the audio content in the audio file is "Open the portal", and the corresponding audio text generated by recognizing the audio file is "Open the portal".
[0056] S130. From at least one pre-acquired dialogue text to be dubbed, determine the target dialogue text that matches the audio text.
[0057] The text to be dubbed refers to pre-prepared text content. This may include game character dialogue, system prompts, or background story narration. Each text to be dubbed has a unique identifier. The target text is the text selected from the texts to be dubbed that is closest to or completely identical to the spoken text.
[0058] In this embodiment of the invention, the voice text and each dialogue text to be dubbed can be compared to calculate the similarity between the voice text and each dialogue text to be dubbed, and then the target dialogue text can be determined based on the similarity. For example, if the voice text is "Open the portal", then the target dialogue text is the preset "Open the portal" dialogue in the game.
[0059] To enhance the flexibility and accuracy of matching, the target dialogue text can be determined by combining the relevant scene content of the corresponding audio file. For example, in some cases, the audio text may contain vague or incomplete sentences. In such cases, the most suitable dialogue text to be dubbed can be selected as the target dialogue text based on the scene content such as the game scene or character type corresponding to the audio text.
[0060] To improve the matching accuracy between speech and dialogue, at least one text to be dubbed can be obtained during the process of determining the target text to match the speech text from at least one pre-acquired text to be dubbed; the speech text is then matched with the at least one text to be dubbed to obtain the text matching degree corresponding to the at least one text to be dubbed; and the target text to be dubbed is determined based on the text to be dubbed corresponding to the text with the maximum text matching degree.
[0061] Text matching score is used to characterize the degree of similarity between the spoken text and the text to be dubbed. For example, text matching score can be expressed as a numerical value, percentage, or score. A higher text matching score indicates a greater similarity between the spoken text and the text to be dubbed, while a lower text matching score indicates a less similarity between the spoken text and the text to be dubbed.
[0062] In this embodiment of the invention, at least one text of dialogue to be dubbed can be obtained from a local database or a remote server via an interface. Further, upon determining the audio text of each audio file, a matching algorithm can be used to match the audio text with each text of dialogue to be dubbed, obtaining the text matching degree corresponding to each text of dialogue to be dubbed. The text matching degrees can be sorted, and the text of dialogue to be dubbed corresponding to the highest text matching degree can be taken as the target text of dialogue to be matched with the audio text. Optionally, the matching algorithm includes, but is not limited to: fuzzy matching algorithms, string-based matching algorithms (such as those based on Euclidean distance, Hamming distance, etc.), rule-based matching algorithms (such as regular expressions), natural language processing (NLP) models (such as term frequency-inverse document frequency, word embedding, etc.) or machine learning algorithms (such as support vector machines, random forests, neural networks, etc.).
[0063] It's important to note that among numerous dialogue texts to be dubbed, the same line may be used by multiple characters. For example, in a game scene, multiple virtual characters might utter the exact same lines. In such cases, text content matching alone is insufficient to accurately determine the target dialogue text. Therefore, to further improve the accuracy of matching between voice and dialogue, when different data sources determine the file identifier of the voice file, the corresponding character type can be filled into the file identifier. Then, by combining the character types corresponding to the voice and dialogue, the target dialogue text that matches the voice text can be filtered.
[0064] In this embodiment of the invention, the target dialogue text that matches the voice text can be determined based on the dialogue text to be dubbed corresponding to the maximum text matching degree. This includes: if there are at least two dialogue texts to be dubbed corresponding to the maximum text matching degree, then the text identifier of each dialogue text to be dubbed corresponding to the maximum text matching degree is determined; if the role type in any text identifier is consistent with the role type in the file identifier of the voice file, then the dialogue text to be dubbed corresponding to the consistent text identifier is determined as the target dialogue text.
[0065] The text identifier can be used to uniquely identify the dialogue text to be dubbed. For example, the text identifier can contain metadata such as text ID, character type, scene information, and text content. The character type can be a classification label used to distinguish different characters. For example, character types include, but are not limited to, male, female, and animal.
[0066] Specifically, during the matching process, if multiple dialogue texts with the same maximum text matching score exist, the text identifier of each dialogue text corresponding to the maximum text matching score can be further extracted. The role type in each text identifier is compared with the role type in the file identifier of the audio file. If the role type in the text identifier of a certain dialogue text matches the role type in the file identifier, then the dialogue text corresponding to the matching text identifier can be identified as the target dialogue text.
[0067] The advantage of this approach is that by considering the character type metadata when matching lines and voice, it is possible to filter out the target line text that matches the voice file from multiple line texts with the same content but different characters, thereby improving the accuracy and reliability of the matching.
[0068] S130. Based on the text identifier of the target dialogue text, rename the file identifier of the audio file to synchronously output the audio file and the target dialogue text with the same identifier.
[0069] In this embodiment of the invention, the file identifier of the audio file can be renamed to match the text identifier of the target dialogue text. This allows audio files with the same identifier and the target dialogue text to be output synchronously in the game scene.
[0070] It should be noted that when renaming the file identifiers of audio files based on the text identifiers of the target dialogue text, duplicate names may occur. To ensure the accuracy of the correspondence between each audio file and the dialogue text, a renaming prompt can be generated when any renamed audio file's file identifier is detected to be identical to its text identifier. This prompt allows the user to determine the desired naming method. The user then receives the desired naming method and renames the audio file's file identifier accordingly.
[0071] The naming methods to be used include custom naming methods and repetitive naming methods. Custom naming methods can use file names defined by the user. Repetitive naming methods can allow duplicate filenames.
[0072] Specifically, when renaming the file identifier of an audio file based on the text identifier of the target dialogue text, if it detects that the file identifier of any already renamed audio file is identical to the text identifier of the target dialogue text, a renaming prompt message is generated. This prompt message can be sent to the user's terminal device. The user can receive the prompt message on their terminal device and then determine whether duplicate naming is allowed. If allowed, the system determines that the desired naming method is duplicate naming. When the system receives a duplicate naming request, it renames the audio file's file identifier based on the text identifier. If duplicate naming is not allowed, a custom filename can be entered. Upon receiving a custom filename, the system considers it to have received the desired naming method and then renames the audio file's file identifier based on the custom filename.
[0073] The technical solution provided by this invention involves acquiring at least one audio file; processing each audio file to obtain corresponding audio text; determining a target dialogue text matching the audio text from at least one pre-acquired dialogue text to be dubbed; and renaming the file identifier of the audio file based on the text identifier of the target dialogue text to synchronously output audio files and target dialogue text with the same identifier. This solves the problem of high cost and easy matching errors in the existing technology, which relies on manual renaming, leading to mismatches between audio and dialogue. The solution achieves the acquisition of audio text corresponding to each audio file through recognition. Matching the acquired at least one dialogue text to be dubbed and the audio text to determine the target dialogue text improves the efficiency and accuracy of audio-text matching. Furthermore, renaming the file identifier of the audio file based on the text identifier of the target dialogue text enhances the efficiency and accuracy of matching audio files with dialogue text, reduces costs, and ensures the matching accuracy between synchronously output audio files with the same identifier and target dialogue text, thereby improving the readability and reliability of information obtained by users.
[0074] Example 2
[0075] Figure 2 This is a flowchart of a file processing method according to Embodiment 2 of the present invention. Based on the foregoing embodiments, a lip-sync file corresponding to the speech file can also be determined to synchronously output speech files and lip-sync files with the same identifier. Specific implementation details can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0076] like Figure 2 As shown, the method specifically includes the following steps:
[0077] S210. Perform kana conversion on the audio text of the audio file to obtain the phonetic text.
[0078] In this context, phonetic text can be text that has undergone kana conversion of spoken text, used to represent the pronunciation information of the spoken text. It should be noted that the phonetic characters in the phonetic text can use letters or symbols to record the phonetic units in the language, and these phonetic units can be phonemes, syllables, or other phonetic components.
[0079] In this embodiment, to further enhance the accuracy of the virtual character's lip movements in reproducing spoken text, kana conversion technology can be used to convert the spoken text of the audio file into kana, resulting in phonetic text. That is, the spoken text is converted into phonetic text that directly reflects the pronunciation of the spoken text.
[0080] To improve the efficiency of kana conversion, contextual information from the spoken text can be incorporated to generate phonetic text. For example, in some cases, the same Chinese character may have different pronunciations in different contexts. In such cases, the most appropriate pronunciation can be selected by analyzing the grammatical structure and vocabulary collocation of the current spoken text to obtain the phonetic text.
[0081] Furthermore, the technical solution provided in this embodiment can also support kana conversion operations in multilingual environments. For example, kana conversion processing includes, but is not limited to, conversion from Chinese to Pinyin, from Japanese to Hiragana or Katakana, and from Korean to Romanization. In other words, the technical solution provided in this embodiment can be used to generate lip-sync files corresponding to audio files in different language versions.
[0082] S220: Generate lip-sync files based on speech files and phonetic text.
[0083] Lip-sync files are data files used to guide the lip movements of virtual or animated characters. For example, a lip-sync file contains a series of timestamps and corresponding lip-sync labels (such as A, E, I, O, U, etc.), indicating the lip movements a character should make at specific points in time. Alternatively, a lip-sync file may contain a series of characters (such as characters in spoken or spoken text) and corresponding lip-sync labels (such as A, E, I, O, U, etc.), indicating the lip movements a character should make for specific characters. Lip-sync labels can be used to identify the lip states corresponding to different pronunciations, such as closed mouth, open mouth, smiling, pouting, rounded lips, etc.
[0084] In this embodiment, to ensure consistency between the lip movements of the virtual character corresponding to the speech file and the speech content, linguistic features (such as phonemes) corresponding to each character in the phonetic text can be analyzed based on predefined pronunciation rules or machine learning models, and corresponding lip movement labels can be assigned to them according to the linguistic features. Based on pronunciation rules or machine learning models, the pronunciation features of each syllable in the speech file can be analyzed, and corresponding lip movement labels can be assigned to them according to the pronunciation features. Furthermore, the lip movement labels corresponding to each linguistic feature in the phonetic text and the lip movement labels corresponding to each pronunciation feature in the speech file can be combined to determine the lip movement file. Optionally, the pronunciation rules may include the mapping relationship between different pronunciation features and different lip movement labels.
[0085] Alternatively: Perform feature analysis on the speech in the speech file, extracting pronunciation features and segmenting the continuous speech stream into a series of independent phonemes. Based on the pronunciation features and phonemes, extract acoustic features (such as pitch, speech rate, energy, etc.). Extract features from each syllable in the phonetic text to obtain the language features of the phonetic text. Based on the acoustic features of the speech file and the language features of the phonetic text, generate a lip-sync file.
[0086] It should be noted that the above method of generating lip-sync files is only one possible implementation method in the embodiments of the present invention. Of course, there are other ways to generate lip-sync files based on speech files and phonetic text. Any method that can be used to generate lip-sync files based on speech files and phonetic text falls within the protection scope of the embodiments of the present invention.
[0087] S230. Based on the file identifier of the voice file, name the lip-sync file so as to synchronously output the voice file and the lip-sync file with the same file identifier.
[0088] In this embodiment, the lip-sync file can be named using the file identifier of the audio file, ensuring that the file identifiers of the lip-sync file and the audio file are consistent. This way, when the virtual character speaks the audio text in the audio file, corresponding lip movements will be generated, ensuring consistency between the virtual character's lip movements and the audio content, thus enhancing the realism and immersion of the user experience.
[0089] It should be noted that if the lip-sync file needs to be updated or modified multiple times, a version number can be added to the file identifier to distinguish between different versions of the lip-sync file. When playing a voice file, the latest version of the lip-sync file is located based on the voice file's file identifier to ensure that the virtual character's lip movements change in sync with the voice content.
[0090] The technical solution provided in this embodiment converts the speech text of a speech file into phonetic text, and then generates a lip-sync file based on the speech file and the phonetic text. The lip-sync file is named using the file identifier of the speech file, thus ensuring the accuracy and speed of matching the speech file and the lip-sync file.
[0091] Building upon the aforementioned embodiments, to enable users to clearly understand how to simultaneously output voice files and text to be dubbed with the same identifier, or simultaneously output voice files and lip-sync files with the same identifier, or simultaneously input voice files, text to be dubbed, and lip-sync files with the same identifier, at least two of the voice files, text to be dubbed, and lip-sync files can be applied to the game scene. This allows the virtual character to speak the voice from the voice file while simultaneously controlling the virtual character's lip movements based on the lip-sync file with the same identifier, and displaying the text to be dubbed with the same identifier, thereby improving the readability of the voice in the game scene and enhancing the realism and immersion of the user's gaming experience.
[0092] In this embodiment, the file identifier of the audio file can also be written into the game script. The game script can refer to a script file used to control the game's execution logic; for example, the game script may contain program code that detects game events, controls character behavior, and calls resources.
[0093] Specifically, the file identifier of the voice file can be written into the code or configuration data of the game script that needs to play the voice file. Alternatively, the file identifier can be embedded as a variable, parameter, or resource reference field into the script logic of the script file.
[0094] For example, when voice from a certain voice file is needed for a virtual character's dialogue in the first level, the file identifier of that voice file can be written into the game script corresponding to that plot node. When the voice content changes, only the voice file needs to be replaced and the file identifier in the game script updated, without the need for large-scale modifications to the script logic, thus improving the flexibility and convenience of game development.
[0095] After writing the file identifier into the game script, when the game script is run, the voice file and the text of the dialogue to be dubbed and the lip-sync file with the same identifier as the voice file can be quickly located and loaded based on the file identifier in the game script, so as to achieve the effect of synchronized voice playback with the character's lip movements and dialogue subtitles.
[0096] In this embodiment, in response to the file retrieval event of the written identifier in the game script, the voice file to be output and the output file to be synchronized corresponding to the written identifier are retrieved; when the voice file to be output is played, the output file to be synchronized is output synchronously.
[0097] The "written identifier" refers to a text identifier written into the game script to identify the voice file and its associated lip-sync file and dialogue text. The voice file to be output can be a voice file retrieved based on the written identifier. The files to be synchronized for output include the lip-sync file to be output and / or the dialogue text to be output.
[0098] In actual game scenarios, a file retrieval event with the written identifier in the game script can be detected when the game runs to the code containing the written identifier, or when the game reaches a specific node (such as entering a scene) or level corresponding to the game script. At this time, a voice file with the same written identifier can be retrieved from the location where the voice file is stored. Simultaneously, a lip-sync file with the same written identifier can be retrieved from the location where the lip-sync file is stored, and / or, the dialogue text with the same written identifier can be retrieved from the location where the dialogue text is stored. This ensures that when playing the voice file to be output, the file to be output is output synchronously, guaranteeing consistency between voice, lip-sync, and subtitles, thus improving the user experience.
[0099] In this embodiment, the method of synchronously outputting the file to be output can be as follows: when the file to be output is a lip-sync file, adjust the lip-sync parameters of the virtual character associated with the game script based on the lip-sync file; when the file to be output is dialogue text, display the dialogue text based on a preset display format.
[0100] Among them, lip-shape parameters refer to control parameters used to describe the lip shape state in the facial animation of a virtual character. For example, lip-shape parameters include, but are not limited to, the degree of mouth opening, lip shape, tongue position, and jaw movement trajectory. Preset display format can refer to predefined styles and position rules used to display dialogue text, such as display position (e.g., bottom of the screen, top of the character's head, center of the dialog box), display style (e.g., font, font size, color, background frame style), and display animation (e.g., fade in / out, slide in, word-by-word display).
[0101] In other words, when the file to be synchronized output is a lip-sync file, the lip-sync parameters of the virtual character associated with the game script can be dynamically adjusted based on the lip-sync tags corresponding to multiple timestamps or multiple characters in the file. Alternatively, the lip-sync file can be input into a facial skeleton or deformation model, allowing the facial skeleton or deformation model to adjust the lip-sync parameters of the virtual character associated with the game script based on the lip-sync file. It should be noted that the lip-sync parameters of the virtual character can be gradually updated during the playback of the voice file, ensuring that the character's lip movements are synchronized with the voice content. When the file to be synchronized output is dialogue text, the dialogue text is displayed on the screen or presented to the player as subtitles according to a preset display format. For example, the preset display format could be a dialog box display, scrolling subtitles at the bottom of the screen, or a speech bubble above the character's head.
[0102] The advantage of this setting is that by dynamically adjusting the lip-sync parameters of the virtual character through lip-sync files, the virtual character can achieve natural and accurate lip-sync changes during voice playback, enhancing the realism and immersion of the voice performance. At the same time, presenting the text content consistent with the voice content to the user in a visual way can enhance the readability and interactivity of the voice performance, making it easier for players to understand the voice content in the game.
[0103] Example 3
[0104] As an optional embodiment of the above embodiments, specific application scenario examples are provided to enable those skilled in the art to further understand the technical solutions of the embodiments of the present invention. Specifically, please refer to the following detailed content.
[0105] In this embodiment, all audio files can be prepared in advance; the game dialogue list, which includes at least one dialogue text to be dubbed, can be read; audio files can be batch-recognized to obtain the audio text corresponding to each audio file; the recognized audio text can be fuzzily matched with at least one dialogue text to be dubbed in the game dialogue list; and the successfully matched audio files can be renamed to match the dialogue ID (i.e., text identifier) of the target dialogue text in the game dialogue list. During the matching process, if the matching fails or the file is renamed, a recognition failure message can be generated. The recognition failure message can include the reason for the voice recognition failure, such as network error or duplicate naming.
[0106] The following section uses Japanese as an example to introduce how to generate lip-sync files.
[0107] It can perform Japanese speech recognition on all audio files, outputting the file identifier (such as the file name) and audio text of the audio file; it converts Japanese kanji in the audio file into phonetic text through the kana conversion module; and it generates a lip-sync file based on the audio file and the phonetic text.
[0108] The technical solution of this embodiment achieves a one-to-one correspondence between voice and dialogue by renaming the voice file based on the file identifier of the target dialogue text that matches the voice text. This improves the efficiency and accuracy of the one-to-one correspondence between voice files and dialogue text, ensuring the matching accuracy between voice and dialogue. Simultaneously, after converting the voice file to voice text, an automatic Chinese character-to-Kana conversion function is used to obtain phonetic text. Combining the voice file and phonetic text, a lip-sync file is generated, improving the efficiency and accuracy of lip-sync file generation. Furthermore, based on the file identifier of the voice file, the lip-sync file is named so that while the voice file is playing, the lip-sync parameters of the virtual character are dynamically adjusted using lip-sync files with the same identifier. This achieves natural and accurate lip-sync changes for the virtual character during voice playback, enhancing the user's gaming experience.
[0109] Example 4
[0110] Figure 3 This is a schematic diagram of the structure of a document processing device according to Embodiment 4 of the present invention. Figure 3 As shown, the device includes: a voice file acquisition module 310, a target dialogue text determination module 320, and a renaming module 330.
[0111] The audio file acquisition module 310 is used to acquire at least one audio file;
[0112] The speech-to-text determination module is used to process each speech file to obtain the speech-to-text corresponding to the speech file.
[0113] The target dialogue text determination module 320 is used to determine the target dialogue text that matches the voice text from at least one pre-acquired dialogue text to be dubbed;
[0114] The renaming module 330 is used to rename the file identifier of the audio file based on the text identifier of the target dialogue text, so as to synchronously output the audio file and the target dialogue text with the same identifier.
[0115] The technical solution of this embodiment involves acquiring at least one audio file; processing each audio file to obtain corresponding audio text; determining a target dialogue text that matches the audio text from at least one pre-acquired dialogue text to be dubbed; and renaming the file identifier of the audio file based on the text identifier of the target dialogue text to synchronously output audio files and target dialogue text with the same identifier. This solves the problem of existing technologies relying on manual renaming, which is costly and prone to comparison errors, leading to mismatches between audio and dialogue. It achieves the goal of obtaining audio text corresponding to each audio file by recognizing each audio file. Matching the acquired at least one dialogue text to be dubbed and the audio text to determine the target dialogue text that matches the audio text improves the efficiency and accuracy of audio-text matching. Furthermore, renaming the file identifier of the audio file based on the text identifier of the target dialogue text improves the efficiency and accuracy of matching audio files with dialogue text, reduces costs, and ensures the matching accuracy between synchronously output audio files with the same identifier and target dialogue text, thereby improving the readability and reliability of information obtained by users.
[0116] Based on the above-mentioned device, optionally, the at least one voice file includes voice files determined by different data sources; the file identification naming methods of different data sources are different;
[0117] The data source includes at least a voice input device, a voice generation model, and a voice synthesizer.
[0118] Based on the above-described device, optionally, the voice file is determined based on at least one of the following methods:
[0119] Based on the voice input device, at least one first voice segment is acquired to obtain a voice file, and the voice file is named;
[0120] At least one speech file is generated based on the speech generation model, and the speech file is named.
[0121] At least one second speech segment is synthesized using a speech synthesizer to obtain a speech file, and the speech file is named.
[0122] Based on the above-mentioned device, optionally, the target dialogue text determination module 320 includes:
[0123] The unit for obtaining dialogue text to be dubbed is used to obtain at least one dialogue text to be dubbed.
[0124] The text matching degree determination unit is used to match the speech text with at least one of the dialogue texts to be dubbed, and obtain the text matching degree corresponding to at least one of the dialogue texts to be dubbed;
[0125] The target dialogue text determination unit is used to determine the target dialogue text that matches the voice text based on the dialogue text to be dubbed corresponding to the maximum text matching degree.
[0126] Based on the above-mentioned device, optionally, the target dialogue text determination unit is specifically used to determine the text identifier of each of the dialogue texts to be dubbed corresponding to the maximum text matching degree if the dialogue texts to be dubbed corresponding to the maximum text matching degree include at least two; if the role type in any of the text identifiers is consistent with the role type in the file identifier of the audio file, then the dialogue text to be dubbed corresponding to the consistent text identifier is determined as the target dialogue text.
[0127] Optionally, based on the above-described apparatus, the apparatus may further include:
[0128] The renaming prompt information determination unit is used to generate renaming prompt information when it detects that the file identifier of any renamed voice file is the same as the text identifier, so that the user can determine the naming method to be used based on the renaming prompt information; wherein, the naming method to be used includes a custom naming method or a repetitive naming method;
[0129] The naming method processing unit is used to receive the naming method to be used and rename the file identifier of the voice file based on the naming method to be used.
[0130] Optionally, based on the above-described apparatus, the apparatus may further include:
[0131] The phonetic text determination unit is used to perform kana conversion processing on the phonetic text of the audio file to obtain the phonetic text.
[0132] The lip-sync file determination unit is used to generate a lip-sync file based on the speech file and the phonetic text;
[0133] The lip-sync file naming unit is used to name the lip-sync file based on the file identifier of the audio file, so as to synchronously output the audio file and the lip-sync file with the same file identifier.
[0134] Optionally, based on the above-described apparatus, the apparatus may further include:
[0135] The identifier writing unit is used to write the file identifier of the audio file into the game script.
[0136] Optionally, based on the above-described apparatus, the apparatus may further include:
[0137] The file retrieval unit is used to retrieve the voice file to be output and the output file to be synchronized corresponding to the written identifier in the game script in response to the file retrieval event of the written identifier; wherein, the output file to be synchronized includes the lip-sync file to be output and / or the dialogue text to be output.
[0138] The synchronous output unit is used to synchronously output the audio file to be output while playing the audio file to be output.
[0139] Based on the above-mentioned device, optionally, a synchronization output unit package is used to adjust the lip-sync parameters of the virtual character associated with the game script based on the lip-sync file to be output when the file to be synchronized is a lip-sync file to be output; and to display the lip-sync text to be output based on a preset display format when the file to be synchronized is a dialogue text to be output.
[0140] The file processing apparatus provided in the embodiments of the present invention can execute the file processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.
[0141] Example 5
[0142] Figure 4 This is a schematic diagram of the structure of an electronic device implementing the document processing method of an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0143] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory 12 or a random access memory 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 12 or a computer program loaded from storage unit 18 into the random access memory 13. The random access memory 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, read-only memory 12, and random access memory 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0144] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0145] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as file processing methods.
[0146] In some embodiments, the file processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via read-only memory 12 and / or communication unit 19. When the computer program is loaded into random access memory 13 and executed by processor 11, one or more steps of the file processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the file processing method by any other suitable means (e.g., by means of firmware).
[0147] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0148] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0149] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0150] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0151] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0152] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0153] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from read-only memory 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of the embodiments of the present invention.
[0154] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the file processing method provided in any embodiment of this invention.
[0155] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0156] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0157] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A file processing method characterized by, The method comprises the following steps: acquiring at least one voice file; for each voice file, processing the voice file to obtain voice text corresponding to the voice file; from at least one pre-acquired voice text to be dubbed, determining a target voice text matched with the voice text; based on the text identifier of the target voice text, renaming the file identifier of the voice file to synchronously output the voice file and the target voice text with the same identifier.
2. The method of claim 1, wherein, The at least one voice file comprises voice files determined by different data sources; the file identifiers of different data sources are named in different ways. The data sources at least include a voice input device, a voice generation model and a voice synthesizer.
3. The method of claim 2, wherein, The voice file is determined based on at least one of the following ways: based on the voice input device, collecting at least one first voice segment to obtain a voice file, and naming the voice file; based on the voice generation model, generating at least one voice file, and naming the voice file; based on the voice synthesizer, synthesizing at least one second voice segment to obtain a voice file, and naming the voice file.
4. The method of claim 1, wherein, The method comprises the following steps: acquiring at least one voice text to be dubbed; matching the voice text with at least one voice text to be dubbed respectively to obtain a text matching degree corresponding to at least one voice text to be dubbed; based on the voice text matched with the voice text, determining a target voice text matched with the voice text.
5. The method of claim 4, wherein, The method comprises the following steps: if the voice text matched with the voice text includes at least two, determining the text identifier of each voice text matched with the voice text; if the character type in any text identifier is consistent with the character type in the file identifier of the voice file, determining the voice text corresponding to the consistent text identifier as the target voice text.
6. The method of claim 1, wherein, In the renaming process of the file identifier of the voice file based on the text identifier of the target voice text, the method further comprises the following steps: when detecting that the file identifier of any voice file after the renaming process is the same as the text identifier, generating a renaming prompt information to enable a user to determine a naming method to be used based on the renaming prompt information; wherein the naming method to be used includes a custom naming method or a repeated naming method; receiving the naming method to be used, and renaming the file identifier of the voice file based on the naming method to be used.
7. The method of claim 1, wherein, The method further comprises the following steps: performing a kana conversion process on the voice text of the voice file to obtain a phonetic text; based on the voice file and the phonetic text, generating a lip file; The mouth shape file is named based on a file identifier of the voice file, so as to synchronously output the voice file and the mouth shape file with the same file identifier.
8. The method according to claim 1 or 7, characterized in that, The method further comprises: writing the file identifier of the voice file into a game script.
9. The method of claim 8, wherein, The method further comprises: in response to a file call event of the written identifier in the game script, calling a to-be-output voice file and a to-be-synchronously-output file corresponding to the written identifier; wherein the to-be-synchronously-output file comprises a to-be-output mouth shape file and / or a to-be-output script text; synchronously outputting the to-be-synchronously-output file while playing the to-be-output voice file.
10. The method of claim 9, wherein, The synchronously outputting the to-be-synchronously-output file comprises: when the to-be-synchronously-output file is a to-be-output mouth shape file, adjusting a mouth shape parameter of a virtual character associated with the game script based on the to-be-output mouth shape file; when the to-be-synchronously-output file is a to-be-output script text, displaying the to-be-output script text based on a preset display form.
11. A file processing apparatus characterized by comprising: comprise: a voice file acquisition module configured to acquire at least one voice file; a voice text determination module configured to, for each voice file, process the voice file to obtain a voice text corresponding to the voice file; a target script text determination module configured to determine, from at least one to-be-voiced script text acquired in advance, a target script text matched with the voice text; a renaming module configured to, based on a text identifier of the target script text, rename a file identifier of the voice file, so as to synchronously output the voice file and the target script text with the same identifier.
12. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the text processing method in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the text processing method in any one of claims 1-10 when executed. The computer readable storage medium stores computer instructions for enabling the processor to implement the text processing method in any one of claims 1-10 when executed.