Cantonese episode voice intelligent cloning and recommending method

By using artificial intelligence and large-scale model analysis to analyze Cantonese opera data, the problem of difficulty in presenting the regional characteristics of Cantonese opera data has been solved, enabling intelligent recommendations and visual analysis for users in different regions, and promoting the research and dissemination of Cantonese opera art.

CN121838732APending Publication Date: 2026-04-10GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610043824.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-13
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively analyze and present the regional characteristics of Cantonese opera data, resulting in an inability to reflect its geographical relevance and hindering the research and promotion of Cantonese opera art.

Method used

By using artificial intelligence and big data models to analyze Cantonese opera data, and through speech synthesis, role matching, keyword extraction and similarity calculation, a visual recommendation list is generated for users in different regions.

Benefits of technology

It enables regional feature analysis and intelligent recommendation of Cantonese opera data, helping users better understand the regional cultural value and promoting the research and dissemination of Cantonese opera art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838732A_ABST
    Figure CN121838732A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a Cantonese episode voice intelligent cloning and recommendation method, which comprises the following steps: firstly, obtaining a voice synthesis text, interaction content of a user and a large language model and voice uploaded by the user; inputting the speech synthesis text into a pre-trained model to generate speech with Cantonese episode characteristics; system cue words are determined through role matching based on user interaction content so as to help the large language model to locate the identity; extracting display keywords in the user interaction content, and analyzing and extracting potential keywords through a large language model; calculating the voice similarity between the voice uploaded by the user and the reference audio in the voice sample database; generating a reference audio recommendation list based on a matching result of three dimensions of the display keyword, the potential keyword and the voice similarity; and finally, presenting the recommendation list through a map visual interface. According to the method, the large model is adopted to analyze Cantonese drama data, regional culture features are visually presented, and user needs are effectively identified and intelligent recommendation is performed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a method for intelligent cloning and recommending Cantonese opera voice recordings. Background Technology

[0002] Cantonese opera is a traditional local opera sung in the Cantonese dialect, possessing unique regional singing characteristics. With the development of artificial intelligence technology, many scholars have conducted research on the digitization of Cantonese opera, including the classification of Cantonese opera schools and singing styles, audio restoration, and text-to-video generation. However, research on Cantonese opera voice cloning is relatively limited. Furthermore, under the influence of different regional cultures (such as Guangzhou, Eastern Guangdong, and Northern Guangdong), Cantonese opera has developed distinct singing styles. Current common presentation methods, such as list formats or streaming media, are insufficient for effectively analyzing and presenting Cantonese opera data, failing to reflect its geographical relevance and other characteristics, thus hindering the research and promotion of Cantonese opera art.

[0003] In view of this, the present invention provides a method for intelligent cloning and recommendation of Cantonese opera voice recordings, which uses artificial intelligence and large models to analyze Cantonese opera data in order to achieve applications such as visual analysis and intelligent recommendation for users in different regions. Summary of the Invention

[0004] To address the aforementioned issues, this invention provides a method for intelligent cloning and recommendation of Cantonese opera voice recordings. It utilizes artificial intelligence and large-scale models to analyze Cantonese opera data, enabling the application of technologies such as visual analysis and intelligent recommendation for users in different regions.

[0005] The above-mentioned technical objective of the present invention is achieved through the following technical solution: A method for intelligent cloning and recommendation of Cantonese opera voice recordings includes the following steps: S1. Obtain the speech-synthesized text, the user's interaction with the large language model, and the user's uploaded speech; S2. Input the synthesized text into the pre-trained GPT-SoVITS model to generate Cantonese opera speech; S3. Based on the interaction content, perform role matching to determine the system prompt words; S4. Extract the display keywords from the interactive content, and extract the potential keywords using the large language model; S5. Calculate the audio similarity between the user-uploaded voice and the reference audio stored in the voice sample database; S6. Based on the matching results of the displayed keywords, potential keywords, and voice similarity, generate a list of recommended reference audio; S7. Visualize the list of recommended reference audios.

[0006] As a preferred method, determining system prompt words through role matching includes the following steps: A database of Cantonese opera character characteristics is constructed, which stores the vocal characteristics, cultural background, and language style data of typical Cantonese opera characters. The user interaction content is semantically matched with the Cantonese opera character feature database to calculate character similarity. If the similarity between the characters is greater than a first preset threshold, the matched character features are used as system prompt words, and the system prompt words are input into the large language model to optimize identity positioning.

[0007] As a preferred method, the extraction and display of keywords includes the following steps: Keyword extraction is performed on the user interaction content to identify explicit words directly related to Cantonese opera styles, regions, or singing styles; If the explicit words are identified, they are marked as display keywords; The extraction of potential keywords includes: The semantic expansion analysis of the user interaction content is performed using the large language model to identify Cantonese opera cultural elements implicitly related to the user's intent. If the confidence level of the semantic expansion analysis is greater than the second preset threshold, it is marked as a potential keyword.

[0008] As a preferred method, the calculation of speech similarity includes the following steps: Mel spectrum feature extraction is performed on the user-uploaded voice and the reference audio; Calculate the cosine similarity value of the Mel spectrum features; If the cosine similarity value is greater than the third preset threshold, it is determined to be a high similarity match; The reference audio recommendation list generated based on the matching results of three dimensions includes: The matching results of the displayed keywords, potential keywords, and voice similarity are sorted. The recommended list of reference audio is generated based on the sorting results.

[0009] As a preferred approach, in the semantic expansion analysis, if the user interaction content is detected to contain regional descriptive words, then the association expansion is performed based on the preset Cantonese opera regional cultural knowledge graph; The Cantonese opera regional culture knowledge graph stores data on typical Cantonese opera schools, representative singing styles, and historical activities in various cities within a preset area. If the regional descriptive terms match the nodes in the Cantonese opera regional culture knowledge graph to a degree greater than the fourth preset threshold, then the matched genre characteristics and activity information will be used as supplementary data for potential keywords.

[0010] As a preferred method, the map visualization interface presentation includes the following steps: The geographical locations of the reference audio collections were marked on the electronic map, and the distribution density of Cantonese opera cultural activities was displayed in the form of a heat map. The reference audio recommendation list is dynamically updated in response to the user's location filtering operation; If the user selects a specific geographic location, the background information of the Cantonese opera genre associated with that location and a preview of the recommended audio will be loaded. The pre-trained GPT-SoVITS model is trained through the following steps: Obtain Cantonese opera voice training samples, and remove background music and noise from the Cantonese opera voice training samples; The processed Cantonese opera speech training samples are input into the GPT-SoVITS model for fine-tuning training to obtain the pre-trained GPT-SoVITS model.

[0011] Compared with the prior art, the present invention has the following advantages: 1. This invention trains a GPT-SoVITS model by removing background music and noise from Cantonese opera speech, then uses this model to analyze Cantonese opera data, uses reference audio with Cantonese opera characteristics to infer the language with Cantonese opera characteristics from the input text, and then judges its regional cultural relevance, so as to realize applications such as visualization analysis and intelligent recommendation for users in different regions.

[0012] 2. This invention performs intelligent audio recommendation based on the interaction content between users and a large language model. It obtains system prompt words through role matching to help the large language model locate its own identity. It also proposes an intelligent recommendation algorithm based on three dimensions: displayed keywords, potential keywords, and voice similarity. This algorithm can effectively identify users' direct and potential needs and recommend corresponding reference languages.

[0013] 3. This invention provides users with a new perspective for selecting reference voices or related activities from a geographic information standpoint through map visualization, making it easier for users to connect with the region and related cultural values ​​behind them. Attached Figure Description

[0014] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.

[0015] Figure 1 This is a flowchart illustrating the intelligent cloning and recommendation method for Cantonese opera voice recordings provided in an embodiment of the present invention.

[0016] Figure 2 This is a schematic diagram of the similarity matching process provided in an embodiment of the present invention.

[0017] Figure 3 This is the user interface for interacting with the large model provided in this embodiment of the invention.

[0018] Figure 4 This is a visual presentation interface for the reference language and activity map provided in the embodiments of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Basic Implementation See Figure 1-4 The Cantonese opera voice intelligent cloning and recommendation method provided in this embodiment includes the following steps: Acquire speech-synthesized text, user interactions with the large language model, and user-uploaded speech; The synthesized text is input into a pre-trained GPT-SoVITS model to generate speech with regional culture and Cantonese opera characteristics; Based on the user interaction content, system prompt words are determined through role matching to help the large language model locate its own identity; Extract display keywords from the user interaction content, and analyze the user interaction content using the large language model to extract potential keywords; Calculate the voice similarity between the user-uploaded voice and the reference audio in the voice sample database to determine their regional cultural relevance; Based on the matching results of the three dimensions of displayed keywords, potential keywords, and voice similarity, a list of recommended reference audio is generated; The list of recommended reference audios is presented through a map visualization interface, which supports filtering based on geographical location and displays the geographical distribution characteristics of the reference audios and related cultural activities.

[0021] As a preferred method, determining system prompt words through role matching includes the following steps: A database of Cantonese opera character characteristics is constructed, which stores the vocal characteristics, regional cultural background and language style data of typical Cantonese opera characters; The user interaction content is semantically matched with the Cantonese opera character feature database to calculate character similarity. If the similarity between the characters is greater than a first preset threshold, the matched character features are used as system prompt words, and the system prompt words are input into the large language model to optimize identity positioning.

[0022] As a preferred method, the extraction and display of keywords includes the following steps: Keyword extraction is performed on the user interaction content to identify explicit words that are directly related to Cantonese opera styles, regions, or singing styles and have regional cultural relevance. If the explicit words are identified, they are marked as display keywords; The extraction of potential keywords includes: The semantic expansion analysis of the user interaction content is performed using the large language model to identify Cantonese opera cultural elements and regional cultural relevance that are implicitly related to the user's intent. If the confidence level of the semantic expansion analysis is greater than the second preset threshold, it is marked as a potential keyword.

[0023] As a preferred method, the calculation of speech similarity includes the following steps: Mel spectrum feature extraction is performed on the user-uploaded voice and the reference audio; Calculate the cosine similarity value of the Mel spectrum features; If the cosine similarity value is greater than the third preset threshold, it is determined to be a high similarity match; The reference audio recommendation list generated based on the matching results of three dimensions includes: The matching results of the displayed keywords, potential keywords, and voice similarity are sorted. The recommended list of reference audio is generated based on the sorting results.

[0024] As a preferred approach, in the semantic expansion analysis, if the user interaction content is detected to contain regional descriptive words, then the association expansion is performed based on the preset Cantonese opera regional cultural knowledge graph; The Cantonese Opera regional culture knowledge graph stores data on typical Cantonese Opera schools, representative singing styles, and historical activities in various cities within Guangdong Province. If the regional descriptive terms match the nodes in the Cantonese opera regional culture knowledge graph to a degree greater than the fourth preset threshold, then the matched genre characteristics and activity information will be used as supplementary data for potential keywords.

[0025] As a preferred method, the map visualization interface presentation includes the following steps: The geographical locations of the reference audio collections were marked on the electronic map, and the distribution density of Cantonese opera cultural activities was displayed in the form of a heat map. The reference audio recommendation list is dynamically updated in response to the user's location filtering operation; If the user selects a specific geographic location, the background information of the Cantonese opera genre associated with that location and a preview of the recommended audio will be loaded. The pre-trained GPT-SoVITS model is trained through the following steps: Obtain Cantonese opera voice training samples, and remove background music and noise from the Cantonese opera voice training samples; The processed Cantonese opera speech training samples are input into the GPT-SoVITS model for fine-tuning training to obtain the pre-trained GPT-SoVITS model.

[0026] Example 1 Based on the basic embodiment, this embodiment 1 further provides a more specific method for intelligent cloning and recommending Cantonese opera voice recordings, which includes the following steps: Step S1: Construct a speech sample database. This invention constructs a sample for each artist's voice. Each sample includes approximately 8 seconds of Cantonese opera speech reference, fine-tuned GPT-SoVITS model parameters, recommended parameter tuning, geographic information, logo, author, genre, relevant description, and keywords. Users select samples and use the model parameters in the samples and the reference audio to perform speech synthesis.

[0027] The Cantonese opera audio reference sample refers to the audio obtained after recording audio clips of Cantonese opera performers singing, removing background music, and performing noise reduction.

[0028] The fine-tuned GPT-SoVITS model parameters are obtained by fine-tuning the pre-trained GPT-SoVITS model on a relevant Cantonese opera speech dataset.

[0029] The parameter tuning parameters include temperature, top_P, and top_K.

[0030] The logo is designed based on the characteristics of the sample. The logo can reflect the characteristics of a certain type or style of Cantonese opera speech and is used for visualization on a map.

[0031] Step S2: The large language model used in this embodiment is Deepseek, implemented by calling the corresponding API. The text input to Deepseek is "I want to clone the Cantonese opera Hongpai voice, like the flat throat style of Hong Xiannu's 'Ode to Lychee,' and I need recommendations for suitable reference voices," along with an audio clip. First, role matching is performed. Preset roles include, but are not limited to, Cantonese opera voice recommendation experts and Cantonese opera art experts. Each role corresponds to a prompt. This prompt is used as the system prompt and, along with the user's input to Deepseek, is used as the user prompt. The system prompt is then input into Deepseek to obtain a response and output it.

[0032] The role matching process first uses word segmentation tools such as jieba to break down the data into lexical units ['clone', 'Cantonese opera', 'Hongpai', 'voice', 'Hongxiannu', 'Lychee Song', 'Pinghou style', 'recommend', 'reference', 'voice']. Then, core words matching each role's keywords are extracted (e.g., "Cantonese opera", "Hongpai", "recommend"), and related words are expanded based on an association matrix (e.g., "Hongpai" is associated with "Hongxiannu" and "Pinghou singing style"). Next, a matching score is calculated for each role, considering the keyword matching type: 3 points for a perfect match, 2 points for an extended match, and 1 point for a substring match. This score is then multiplied by the keyword weight to determine the fit between the user's needs and the role. Finally, the roles are sorted by score, and the highest-scoring role is used as the matching result. If the scores are close, a mixed role is returned; if both are 0, the default role is returned.

[0033] Step S3: Select the system role for keyword extraction as the system prompt. Input the above input text and the system prompt together into Deepseek to obtain the genre and keywords used to match the Cantonese opera speech reference samples, for example, ["Cantonese opera", "Hong school", "Hong Xiannu", "Hong style", "female style", "Ode to Lychee", "Chen Guanqing", "Pinghou style", "Pinghou singing", "traditional Cantonese opera", "female singing", "mellow timbre", "classic repertoire", "Cantonese opera singing", "Hong school characteristics", "Hong Xiannu style", "Pinghou singing method", "traditional Cantonese opera", "Lingnan opera", "Cantonese opera vocal style", "Hong school inheritance"]. The obtained genre and keywords will be applied to the intelligent recommendation algorithm below.

[0034] In this embodiment, the system role prompt is described as follows: "You are a professional keyword extraction assistant specializing in Cantonese opera. Your core task is to comprehensively extract and deeply mine directly and potentially related keywords from the input text, including seven categories of information: school-related (including school name, branches, and lineage), representative figures (including school founders and core inheritors), singing styles (including distinctive singing styles, subdivided singing styles, variant singing styles, and vocal techniques), vocal style (including timbre characteristics and style descriptions), classic content (including classic plays and excerpts), and related figures (including playwrights and composers related to the school / singing style). The output only needs to present a list of keywords, without additional text descriptions. List items should be deduplicated and arranged in the logical order of 'school → figure → singing style → style → play → related information.' The length of a single keyword should be controlled between 1 and 6 characters to avoid redundant phrases. Finally, the list should be enclosed in square brackets, with each keyword separated by a comma, for example: ["Cantonese opera", "Hong school", "Hong Xiannu", "Hong style"]. The matching keywords include both the keywords displayed in the user input and the keywords implied behind them.

[0035] Step S4, as shown in Figure 2, performs regular expression matching based on the school and keywords extracted from the user input in step S3, such as Cantonese opera school (Hong School, Ma School, etc.), repertoire type (traditional play, newly adapted play, etc.), singing style (Bangzi, Erhuang, etc.), actor information, duration requirements, etc. Each keyword corresponds to a specific weight, and the corresponding weight set is recorded. ,and ,in For the corresponding number The weight of each keyword, , The total number of keywords. Then calculate the keyword weights matched for each sample. Obtain keyword weights and sets. , The total number of samples is given. Next, implicit matching is performed. The sbert-base-chinese-nli Chinese semantic model is used to convert the user input text and the content in the Cantonese opera voice resource library (including play introductions, singing characteristics, performance styles, etc.) into vectors respectively. Then, the cosine similarity is calculated to obtain the relevance with the samples, which is used as the implicit matching similarity score. The cosine similarity calculation formula is as follows:

[0036] in, and These represent the semantic vectors of the user query and the description of the Cantonese opera voice reference sample, respectively. This indicates the similarity between the two.

[0037] Obtain the cosine similarity set for each sample. Where n is the total number of samples. For the first The cosine similarity is calculated for each sample. .

[0038] Step S5: If the user uploads reference audio, it is matched with the reference language in the samples for similarity, serving as the third recommendation metric. The user-uploaded audio and the audio samples in the language sample database are all sampled at a uniform rate of 16kHz and quantization bit depth of 16bit, converted to mono, and standardized to 8 seconds. If a sample in the language sample database is longer than 8 seconds, the 8-second segment with the highest energy is selected; if it is shorter than 8 seconds, silence is added at the beginning and end to bring it to 8 seconds. The total number of samples is set to 128,000 (8 seconds × 16kHz); the frame length is 400 samples, and the frame shift is 160 samples, resulting in a total of 798 frames. Then, a high-pass filter is used to enhance high-frequency details, as shown in the following formula:

[0039] Then, a 400-sample-point Hamming window is applied to each frame to reduce edge artifacts, as shown in the following formula:

[0040] Where m is the intra-frame sampling index and N is the frame length.

[0041] Perform a 1024-point FFT on the windowed frame, calculate the power spectrum, and then pass it through a 32-mel filter bank, as shown in the following formula:

[0042] Where t is the frame index and m is the filter index. Let be the power spectrum of the t-th frame. For the first The frequency points corresponding to each Mel filter.

[0043] After taking the logarithm of the filtered energy, a DCT is performed to obtain the 13-dimensional MFCC features, i.e., the feature matrix. The MFCC formula is as follows:

[0044] PCA was used to reduce the dimensionality to 8 dimensions, resulting in a matrix. .

[0045] The same method was used to obtain the reference audio uploaded by the user. Calculate the inter-frame distance matrix using Euclidean distance. The calculation formula is as follows: ,

[0046] Where t1 is the user feature frame index, t2 is the sample feature frame index, and k is the feature dimension index after dimensionality reduction.

[0047] Then calculate the DTW accumulated distance C(t) 1, t2), the calculation formula is as follows:

[0048] Boundary conditions are , The smaller the value, the stronger the correlation. and each sample in the speech sample database Calculate DTW accumulated distance ,get , where n is the total number of samples. Find the maximum value. Then normalize the result; the calculation formula is as follows:

[0049] in, , .

[0050] Step S6: Aggregate the above three weight values ​​to obtain the final weight. , The calculation formula is as follows:

[0051] right Sort in descending order. The larger the value, the more relevant it is. Then, the values ​​are presented in a list from largest to smallest relevance, as shown in Figure 3.

[0052] Step S7: Select the 10 most relevant samples and visualize them on a map, as shown in Figure 4. Specifically, call Amap to initialize the map to the location in the sample set, insert the sample's logo at the corresponding position on the map, and embed an introductory link for the sample in the logo. Users can click to view and select the sample. In addition, it provides classification and filtering by characteristics including but not limited to genre, singing style, and author.

[0053] The map visualization method for the reference speech can be applied to the visualization of Cantonese opera performances or other activities, and can be classified and filtered by two map modes: language map mode and activity map mode.

[0054] Step S8: Find the corresponding reference speech, pre-trained GPT-SoVITS, and recommended parameter tuning based on the selected samples. Input the speech text to be cloned and the reference speech. If the user has not uploaded a reference audio, the reference speech from the selected samples will be used. By default, the recommended parameter tuning is used to infer the cloned speech with Cantonese opera characteristics. If the user is not satisfied with the effect, they can adjust the parameters themselves.

[0055] The Cantonese opera voice intelligent cloning and recommendation method provided in this embodiment focuses on using a large model to analyze Cantonese opera data, visually presenting the cultural characteristics of Cantonese opera in different regions, effectively identifying user needs, and making intelligent recommendations, which is beneficial to the research, application, and promotion of Cantonese opera culture.

[0056] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.

[0057] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to achieve the described functions, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the described devices, apparatuses, and units can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0058] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, function, and operation of possible implementations of apparatus, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than those disclosed in the description; sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based device that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Claims

1. A method for intelligent cloning and recommending Cantonese opera voice recordings, characterized in that, Includes the following steps: S1. Obtain the speech-synthesized text, the user's interaction with the large language model, and the user's uploaded speech; S2. Input the synthesized text into the pre-trained GPT-SoVITS model to generate Cantonese opera speech; S3. Based on the interaction content, perform role matching to determine the system prompt words; S4. Extract the display keywords from the interactive content, and extract the potential keywords using the large language model; S5. Calculate the audio similarity between the user-uploaded voice and the reference audio stored in the voice sample database; S6. Based on the matching results of the displayed keywords, potential keywords, and voice similarity, generate a list of recommended reference audio; S7. Visualize the list of recommended reference audios.

2. The method for intelligent cloning and recommending Cantonese opera voice recordings according to claim 1, characterized in that, The process of determining system prompt words through role matching includes the following steps: A database of Cantonese opera character characteristics is constructed, which stores the vocal characteristics, cultural background, and language style data of typical Cantonese opera characters. The user interaction content is semantically matched with the Cantonese opera character feature database to calculate character similarity. If the similarity between the characters is greater than a first preset threshold, the matched character features are used as system prompt words, and the system prompt words are input into the large language model.

3. The method for intelligent cloning and recommending Cantonese opera voices according to claim 1 or 2, characterized in that, The extraction and display of keywords includes the following steps: Keyword extraction is performed on the user interaction content to identify explicit words directly related to Cantonese opera styles, regions, or singing styles; If the explicit words are identified, they are marked as display keywords; The extraction of potential keywords includes: The semantic expansion analysis of the user interaction content is performed using the large language model to identify Cantonese opera cultural elements implicitly related to the user's intent. If the confidence level of the semantic expansion analysis is greater than the second preset threshold, it is marked as a potential keyword.

4. The method for intelligent cloning and recommending Cantonese opera voice recordings according to claim 1 or 2, characterized in that, The calculation of speech similarity includes the following steps: Mel spectrum feature extraction is performed on the user-uploaded voice and the reference audio; Calculate the cosine similarity value of the Mel spectrum features; If the cosine similarity value is greater than the third preset threshold, it is determined to be a high similarity match; The reference audio recommendation list generated based on the matching results of three dimensions includes: The matching results of the displayed keywords, potential keywords, and voice similarity are sorted. The recommended list of reference audio is generated based on the sorting results.

5. The method for intelligent cloning and recommending Cantonese opera voices according to claim 3, characterized in that, In the semantic expansion analysis, if the user interaction content is detected to contain regional descriptive words, then the association expansion is performed based on the preset Cantonese opera regional cultural knowledge graph. The Cantonese opera regional culture knowledge graph stores data on typical Cantonese opera schools, representative singing styles, and historical activities in various cities within a preset area. If the regional descriptive terms match the nodes in the Cantonese opera regional culture knowledge graph to a degree greater than the fourth preset threshold, then the matched genre characteristics and activity information will be used as supplementary data for potential keywords.

6. The method for intelligent cloning and recommending Cantonese opera voice recordings according to claim 5, characterized in that, The map visualization interface is presented through the following steps: The geographical locations of the reference audio collections were marked on the electronic map, and the distribution density of Cantonese opera cultural activities was displayed in the form of a heat map. The reference audio recommendation list is dynamically updated in response to the user's location filtering operation; If the user selects a specific geographic location, the background information of the Cantonese opera genre associated with that location and a preview of the recommended audio will be loaded. The pre-trained GPT-SoVITS model is trained through the following steps: Obtain Cantonese opera voice training samples, and remove background music and noise from the Cantonese opera voice training samples; The processed Cantonese opera speech training samples are input into the GPT-SoVITS model for fine-tuning training to obtain the pre-trained GPT-SoVITS model.