A method and device for generating speech synthesis information

By extracting and integrating human joint nodes and lip characteristics and combining multimodal models for speech synthesis, the problem of communication difficulties for language barriers is solved and the business processing efficiency of the financial service industry is improved.

CN115240637BActive Publication Date: 2025-08-01INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210870161.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-22
Publication Date
2025-08-01
Estimated Expiration
2042-07-22

AI Technical Summary

Technical Problem

In the financial service industry, it is difficult for language-impaired customers to communicate with practitioners, and the existing technology cannot fully obtain information about language-impaired customers, resulting in inefficient business processing.

Method used

By obtaining target videos, extracting human joint node features and lip characteristics, and fusion perform text recognition and speech synthesis, using multimodal models to improve the accuracy of information acquisition, including the recognition of finger joints and target features.

Benefits of technology

It realizes accurate identification and efficient communication of voice information for customers with language barriers, and improves business processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240637B_ABST
    Figure CN115240637B_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for generating speech synthesis information, which relates to the technical field of speech data processing and can be used in the financial field or other technical fields. The method includes: obtaining a target video, and extracting human joint point features and lip language features according to the target video; the target video includes real-time images of customers with speech disorders; fusing the human joint point features and the lip language features to obtain a fused feature, and performing text recognition on the fused feature to obtain text vocabulary information; performing speech synthesis on the text vocabulary information to obtain speech synthesis information. The apparatus executes the above method. The method and apparatus for generating speech synthesis information provided by the embodiments of the present invention can accurately and efficiently identify the speech information that customers with speech disorders want to express, and improve the business handling efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech data processing, and particularly to a method and device for generating speech synthesis information. Background Art

[0002] In the financial service industry, practitioners often need to face customers with language barriers. However, most practitioners do not understand sign language and cannot communicate with such customers normally. For customers with language barriers, they can only communicate with practitioners by writing or using sign language, which causes communication difficulties and leads to problems such as low business handling efficiency.

[0003] The prior art recognizes the sign language of customers with language barriers and then converts it into the speech information that the customers with language barriers want to express. However, the recognized sign language information is single and cannot fully obtain the information that the customers with language barriers want to express. Summary of the Invention

[0004] In view of the problems in the prior art, embodiments of the present invention provide a method and device for generating speech synthesis information, which can at least partially solve the problems existing in the prior art.

[0005] On the one hand, the present invention proposes a method for generating speech synthesis information, including:

[0006] Obtain a target video, and extract human joint point features and lip language features according to the target video; the target video includes real-time images of customers with language barriers;

[0007] Fuse the human joint point features and the lip language features to obtain a fused feature, and perform text recognition on the fused feature to obtain text vocabulary information;

[0008] Perform speech synthesis on the text vocabulary information to obtain speech synthesis information.

[0009] Wherein, the human joint point features include finger joint features; correspondingly, the method for generating speech synthesis information further includes:

[0010] If a target object pointed to by a finger is recognized based on the finger joint features, extract the target object features;

[0011] Fuse the human joint point features, the lip language features and the target object features to obtain the fused feature.

[0012] Wherein, the performing text recognition on the fused feature to obtain text vocabulary information includes:

[0013] Perform text recognition on the fused feature based on a preset feature recognition model to obtain text vocabulary information;

[0014] Among them, the preset feature recognition model is obtained by training a neural network based on feature recognition sample data.

[0015] Among them, before the step of performing text recognition on the fusion feature based on the preset feature recognition model, the voice synthesis information generation method further includes:

[0016] Perform sentence segmentation detection on the target video to obtain sentence segmentation nodes, and segment the target video with the sentence segmentation nodes to obtain each target video segment;

[0017] Perform text recognition on the fusion features in each target video segment respectively based on the preset feature recognition model.

[0018] Among them, the voice synthesis information generation method further includes:

[0019] Fuse the text vocabulary information and the target video to obtain comprehensive fusion information;

[0020] Perform voice synthesis on the comprehensive fusion information to obtain voice synthesis information.

[0021] Among them, the performing voice synthesis on the comprehensive fusion information to obtain voice synthesis information includes:

[0022] Perform voice synthesis on the comprehensive fusion information based on a preset voice synthesis model to obtain voice synthesis information;

[0023] Among them, the preset voice synthesis model is a multi-modal model that can perform voice synthesis on video content and language content respectively.

[0024] Among them, after the step of obtaining the voice synthesis information, the voice synthesis information generation method further includes:

[0025] Output the voice synthesis information through a speaker.

[0026] On the one hand, the present invention proposes a voice synthesis information generation device, including:

[0027] An acquisition unit, configured to acquire a target video, and extract human joint point features and lip language features according to the target video; the target video includes a real-time image of a customer with speech disorders;

[0028] A fusion unit, configured to fuse the human joint point features and the lip language features to obtain a fusion feature, and perform text recognition on the fusion feature to obtain text vocabulary information;

[0029] A synthesis unit, configured to perform voice synthesis on the text vocabulary information to obtain voice synthesis information.

[0030] In another aspect, an embodiment of the present invention provides an electronic device, including: a processor, a memory, and a bus, where,

[0031] The processor and the memory complete communication with each other through the bus;

[0032] The memory stores program instructions executable by the processor, and the processor can execute the following method by invoking the program instructions:

[0033] Obtain a target video, and extract human joint point features and lip language features according to the target video; the target video includes real-time images of customers with speech impairments;

[0034] Fuse the human joint point features and the lip language features to obtain a fused feature, and perform text recognition on the fused feature to obtain text vocabulary information;

[0035] Perform speech synthesis on the text vocabulary information to obtain speech synthesis information.

[0036] An embodiment of the present invention provides a non-transitory computer-readable storage medium, including:

[0037] The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the following method:

[0038] Obtain a target video, and extract human joint point features and lip language features according to the target video; the target video includes real-time images of customers with speech impairments;

[0039] Fuse the human joint point features and the lip language features to obtain a fused feature, and perform text recognition on the fused feature to obtain text vocabulary information;

[0040] Perform speech synthesis on the text vocabulary information to obtain speech synthesis information.

[0041] The speech synthesis information generation method and device provided by the embodiment of the present invention obtain a target video, and extract human joint point features and lip language features according to the target video; the target video includes real-time images of customers with speech impairments; fuse the human joint point features and the lip language features to obtain a fused feature, and perform text recognition on the fused feature to obtain text vocabulary information; perform speech synthesis on the text vocabulary information to obtain speech synthesis information, which can accurately and efficiently identify the speech information that customers with speech impairments want to express, and improve the business handling efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:

[0043] Figure 1 It is a schematic flowchart of a method for generating speech synthesis information provided by an embodiment of the present invention.

[0044] Figure 2 It is a schematic flowchart of a method for generating speech synthesis information provided by another embodiment of the present invention.

[0045] Figure 3 It is a schematic structural diagram of modularization of a method for generating speech synthesis information provided by an embodiment of the present invention.

[0046] Figure 4 It is a schematic structural diagram of a device for generating speech synthesis information provided by an embodiment of the present invention.

[0047] Figure 5 It is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the following will further describe the embodiments of the present invention in detail with reference to the drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention. It should be noted that, without conflict, the embodiments and features in the embodiments in this application can be combined with each other arbitrarily.

[0049] The relevant terms of the embodiments of the present invention are uniformly described as follows:

[0050] 1. Multimodal technology

[0051] It refers to a modeling technology that integrates or fuses two or more information modalities.

[0052] 2. Speech synthesis

[0053] Also known as Text to Speech (TTS) technology, it is a technology that converts the text information generated by a computer itself or input externally into audible and fluent spoken language output.

[0054] 3. Lip reading

[0055] Lip reading, also known as lip language, lip reading or speechreading, is a method of understanding language by visually interpreting the movements of the lips, face, and tongue when normal sounds cannot be heard (e.g., due to hearing impairment or insufficient sound recording).

[0056] 4. Sign language

[0057] A means of communication using finger spelling and gestures, mostly used by deaf-mutes.

[0058] 5. Face recognition

[0059] That is, face detection, which is a computer technology for finding the location and size of a human face in any digital image.

[0060] 6. Human keypoints detection

[0061] Human Keypoints Detection, also known as human pose estimation, is to find the positions of joints including the head, hands, elbows, etc. in the human body, and then connect them in sequence to obtain a "stick figure".

[0062] 7. Video content understanding

[0063] The video content understanding ability can analyze various elements such as stars, ordinary people, and game scenes in the video.

[0064] 8. Natural language model

[0065] The probability distribution of word sequences, which is used to determine a probability distribution for a text of length m, indicating the possibility of the existence of the text. Text error correction can be achieved through a large number of real-world texts.

[0066] Figure 1 It is a schematic flowchart of the method for generating speech synthesis information provided by an embodiment of the present invention. As Figure 1 shown, the method for generating speech synthesis information provided by the embodiment of the present invention includes:

[0067] Step S1: Obtain a target video, and extract human joint point features and lip reading features according to the target video; the target video contains real-time images of customers with language impairments.

[0068] Step S2: Fuse the human joint point features and the lip reading features to obtain a fused feature, and perform text recognition on the fused feature to obtain text vocabulary information.

[0069] Step S3: Perform speech synthesis on the text vocabulary information to obtain speech synthesis information.

[0070] In the above step S1, the device acquires a target video and extracts human joint point features and lip language features from the target video; the target video includes the real-time image of a customer with a language barrier. The device may be a computer device that executes this method, such as a server. It should be noted that the acquisition and analysis of data involved in the embodiments of the present invention are authorized by the user.

[0071] The target video can be acquired by a camera, and the target video is obtained by receiving the video stream of the camera.

[0072] Customers with language barriers may include customers who are unable to speak or have difficulty expressing themselves and come to the bank counter to handle business.

[0073] The human joint point features are the position features of joints including the head, hands, elbows, etc. in the above human body. Since the customer needs to sit down to handle business, therefore, the human joint point features mainly refer to the position features of the joints of the human body parts above the chest. By identifying the human joint point features, the human body posture can be determined, and the voice information that the customer wants to express can be determined according to the human body posture.

[0074] The lip language features are the features that reflect the characteristics of lip language actions, and can be specifically referred to the above description.

[0075] The human joint point features can be extracted by human key point detection technology. The extraction of lip language features can also be realized by using existing mature technologies.

[0076] In the above step S2, the device fuses the human joint point features and the lip language features to obtain a fusion feature, and performs text recognition on the fusion feature to obtain text vocabulary information. At this time, the fusion feature includes human joint point features and lip language features, and thus the information volume of the voice synthesis information includes human joint point features and lip language features.

[0077] The human joint point features include finger joint features; correspondingly, the method for generating the voice synthesis information further includes:

[0078] If the target object pointed to by the finger is recognized based on the finger joint features, the target object features are extracted; by recognizing the finger joint features, it can be recognized where the user's finger points. If it is determined that there is a target object pointed to by the finger, the target object features are extracted. The target object can be an item carried by the customer, such as an ID card, a mobile phone, a business handling document, etc.; it can also be a related item near the bank counter, such as a password input device, a signature pen, etc.

[0079] By extracting the target object features, it can be determined what kind of target object the customer's finger points to.

[0080] Fuse the human joint point features, the lip-reading features, and the target object features to obtain the fused features. The fused features at this time include human joint point features, lip-reading features, and target object features. Furthermore, the amount of information of the speech synthesis information obtained includes human joint point features, lip-reading features, and target object features.

[0081] Performing text recognition on the fused features to obtain text vocabulary information, including:

[0082] Performing text recognition on the fused features based on a preset feature recognition model to obtain text vocabulary information; the fused features can be vectorized and input into the preset feature recognition model, and the output result of the preset feature recognition model is used as the text vocabulary information. The text vocabulary information is the information that the customer wants to express in the form of text vocabulary.

[0083] Among them, the preset feature recognition model is obtained by training a neural network according to feature recognition sample data.

[0084] The feature recognition sample data can include human joint point sample data, lip-reading sample data, and target object sample data. The specific type of the neural network and the training method can adopt existing methods and will not be elaborated here.

[0085] Before the step of performing text recognition on the fused features based on the preset feature recognition model, the speech synthesis information generation method further includes:

[0086] Performing sentence segmentation detection on the target video to obtain sentence segmentation nodes, and segmenting the target video with the sentence segmentation nodes to obtain each target video segment; each frame image in the target video can be detected, and it is sequentially determined whether each frame image meets the sentence segmentation condition. If a frame image that meets the sentence segmentation condition is detected, it can be determined that there is a sentence segmentation node. The specific position of the sentence segmentation node can be realized through a classification network or by detecting the time point of a long sign language-free action through joint point detection.

[0087] Performing text recognition on the fused features in each target video segment based on the preset feature recognition model respectively. The description of performing text recognition on the fused features in the overall target video above can be referred to and will not be elaborated here.

[0088] In the above step S3, the device performs speech synthesis on the text vocabulary information to obtain speech synthesis information. The TTS technology can be used to perform speech synthesis on the text vocabulary information to obtain speech synthesis information. As Figure 2 shown, the speech synthesis information generation method further includes:

[0089] Fuse the text vocabulary information and the target video to obtain comprehensive fusion information; fuse the text vocabulary information output by the recognition network and the target video captured by the camera to obtain comprehensive fusion information.

[0090] Perform speech synthesis on the comprehensive fusion information to obtain speech synthesis information. The TTS technology can be used to perform speech synthesis on the comprehensive fusion information to obtain speech synthesis information.

[0091] The step of performing speech synthesis on the comprehensive fusion information to obtain speech synthesis information includes:

[0092] Perform speech synthesis on the comprehensive fusion information based on a preset speech synthesis model to obtain speech synthesis information;

[0093] Among them, the preset speech synthesis model is a multi-modal model that can perform speech synthesis on video content and language content respectively. As Figure 2 shown, the natural language understanding and video content understanding capabilities of the multi-modal model are used to associate the customer's semantics with the video content, more accurately reflecting the information that the customer wants to express.

[0094] The advantage of using the multi-modal model is that it can complete accurate content understanding with less data.

[0095] As Figure 2 shown, after the step of obtaining the speech synthesis information, the method for generating the speech synthesis information further includes:

[0096] Output the speech synthesis information through a speaker.

[0097] As Figure 3 shown, the method of the embodiment of the present invention can be implemented based on modularization, which is described as follows:

[0098] It is mainly divided into four major modules: data acquisition, feature extraction, vocabulary recognition, and semantic expression. Among them, the data acquisition module mainly completes through a camera to capture video images related to the target person.

[0099] After obtaining the video image information, a part is sent to the feature extraction module. The human joint point features are extracted through the human key point detection technology, the lip language features are extracted through the lip feature module, and the two features are fused and then sent to the recognition network.

[0100] The recognition network includes a sentence-breaking and recognition function. When there is no vocabulary expression, the sentence-breaking function outputs empty, and when there is a vocabulary expression, it outputs the corresponding vocabulary. The vocabulary sequence is sent to the semantic expression module.

[0101] The semantic expression module integrates video image information and recognized vocabulary information, and uses the natural language understanding and video content understanding capabilities of the multimodal model to associate the user's semantics with the video content, more accurately reflecting the information that the target person wants to express. The advantage of using the multimodal model is that it can complete accurate content understanding with less data. The text information output by the semantic expression module is sent to the speech synthesis module of the multimodal model. After completing speech generation, it is output to the speaker to obtain the final audio output.

[0102] The speech synthesis information generation method provided by the embodiments of the present invention obtains a target video, and extracts human joint point features and lip language features according to the target video; the target video includes real-time images of customers with speech disorders; the human joint point features and the lip language features are fused to obtain a fused feature, and text recognition is performed on the fused feature to obtain text vocabulary information; speech synthesis is performed on the text vocabulary information to obtain speech synthesis information, which can accurately and efficiently recognize the speech information that customers with speech disorders want to express, and improve the business handling efficiency.

[0103] Further, the human joint point features include finger joint features; correspondingly, the speech synthesis information generation method further includes:

[0104] If the target object pointed to by the finger is recognized based on the finger joint features, the target object features are extracted; reference may be made to the above description and will not be repeated.

[0105] The human joint point features, the lip language features and the target object features are fused to obtain the fused feature. Reference may be made to the above description and will not be repeated.

[0106] The speech synthesis information generation method provided by the embodiments of the present invention can further accurately and efficiently recognize the speech information that customers with speech disorders want to express and improve the business handling efficiency by fusing the target object features.

[0107] Further, the performing text recognition on the fused feature to obtain text vocabulary information includes:

[0108] Performing text recognition on the fused feature based on a preset feature recognition model to obtain text vocabulary information; reference may be made to the above description and will not be repeated.

[0109] Wherein, the preset feature recognition model is obtained by training a neural network according to feature recognition sample data. Reference may be made to the above description and will not be repeated.

[0110] The speech synthesis information generation method provided by the embodiments of the present invention can accurately and conveniently obtain text vocabulary information.

[0111] Further, before the step of performing text recognition on the fusion feature based on the preset feature recognition model, the method for generating speech synthesis information further includes:

[0112] Perform sentence segmentation detection on the target video to obtain sentence segmentation nodes, and segment the target video with the sentence segmentation nodes to obtain each target video segment; reference may be made to the above description and will not be elaborated herein.

[0113] Perform text recognition on the fusion features in each target video segment respectively based on the preset feature recognition model. Reference may be made to the above description and will not be elaborated herein.

[0114] The method for generating speech synthesis information provided by the embodiment of the present invention can further accurately and conveniently obtain text vocabulary information by performing text recognition on the fusion features in segments.

[0115] Further, the method for generating speech synthesis information further includes:

[0116] Fuse the text vocabulary information and the target video to obtain comprehensive fusion information; reference may be made to the above description and will not be elaborated herein.

[0117] Perform speech synthesis on the comprehensive fusion information to obtain speech synthesis information. Reference may be made to the above description and will not be elaborated herein.

[0118] The method for generating speech synthesis information provided by the embodiment of the present invention can further accurately and efficiently recognize the speech information that a customer with a language disorder wants to express by performing speech synthesis on the comprehensive fusion information, improving the business handling efficiency.

[0119] Further, the step of performing speech synthesis on the comprehensive fusion information to obtain speech synthesis information includes:

[0120] Perform speech synthesis on the comprehensive fusion information based on a preset speech synthesis model to obtain speech synthesis information; reference may be made to the above description and will not be elaborated herein.

[0121] Wherein, the preset speech synthesis model is a multimodal model that can perform speech synthesis on video content and language content respectively. Reference may be made to the above description and will not be elaborated herein.

[0122] The method for generating speech synthesis information provided by the embodiment of the present invention can improve the accuracy and processing efficiency of speech synthesis by performing speech synthesis on the comprehensive fusion information through the model.

[0123] Further, after the step of obtaining the speech synthesis information, the method for generating speech synthesis information further includes:

[0124] Output the speech synthesis information through a speaker. Reference may be made to the above description and will not be elaborated herein.

[0125] The voice synthesis information generation method provided by the embodiment of the present invention enables bank staff and customers to hear voice information in a timely manner, further facilitating business handling.

[0126] It should be noted that the voice synthesis information generation method provided by the embodiment of the present invention can be used in the financial field and can also be used in any technical field other than the financial field. The embodiment of the present invention does not limit the application field of the voice synthesis information generation method.

[0127] Figure 4 It is a structural schematic diagram of a voice synthesis information generation device provided by an embodiment of the present invention. As Figure 4 shown, the voice synthesis information generation device provided by the embodiment of the present invention includes an acquisition unit 401, a fusion unit 402, and a synthesis unit 403, where:

[0128] The acquisition unit 401 is used to acquire a target video and extract human joint point features and lip language features according to the target video; the target video includes real-time images of customers with speech impairments; the fusion unit 402 is used to fuse the human joint point features and the lip language features to obtain a fusion feature, and perform text recognition on the fusion feature to obtain text vocabulary information; the synthesis unit 403 is used to perform voice synthesis on the text vocabulary information to obtain voice synthesis information.

[0129] Specifically, the acquisition unit 401 in the device is used to acquire a target video and extract human joint point features and lip language features according to the target video; the target video includes real-time images of customers with speech impairments; the fusion unit 402 is used to fuse the human joint point features and the lip language features to obtain a fusion feature, and perform text recognition on the fusion feature to obtain text vocabulary information; the synthesis unit 403 is used to perform voice synthesis on the text vocabulary information to obtain voice synthesis information.

[0130] The voice synthesis information generation device provided by the embodiment of the present invention acquires a target video, extracts human joint point features and lip language features according to the target video; the target video includes real-time images of customers with speech impairments; fuses the human joint point features and the lip language features to obtain a fusion feature, and performs text recognition on the fusion feature to obtain text vocabulary information; performs voice synthesis on the text vocabulary information to obtain voice synthesis information, which can accurately and efficiently identify the voice information that customers with speech impairments want to express and improve the business handling efficiency.

[0131] Further, the human joint point features include finger joint features; correspondingly, the voice synthesis information generation device is further used for:

[0132] If the target object pointed by the finger is recognized based on the finger joint features, the target object features are extracted;

[0133] Fuse the human joint point features, the lip language features and the target object features to obtain the fused features.

[0134] The speech synthesis information generation device provided by the embodiment of the present invention can further accurately and efficiently recognize the speech information that a language-impaired customer wants to express by fusing the target object features, improving the business handling efficiency.

[0135] Further, the fusion unit 402 is specifically configured to:

[0136] Perform text recognition on the fused features based on a preset feature recognition model to obtain text vocabulary information;

[0137] Wherein, the preset feature recognition model is obtained by training a neural network according to feature recognition sample data.

[0138] The speech synthesis information generation device provided by the embodiment of the present invention can accurately and conveniently obtain text vocabulary information.

[0139] Further, before the step of performing text recognition on the fused features based on the preset feature recognition model, the speech synthesis information generation device is specifically configured to:

[0140] Perform sentence segmentation detection on the target video to obtain sentence segmentation nodes, and segment the target video with the sentence segmentation nodes to obtain each target video segment;

[0141] Perform text recognition on the fused features in each target video segment respectively based on the preset feature recognition model.

[0142] The speech synthesis information generation device provided by the embodiment of the present invention can further accurately and conveniently obtain text vocabulary information by performing text recognition on the fused features in segments.

[0143] Further, the speech synthesis information generation device is specifically configured to:

[0144] Fuse the text vocabulary information and the target video to obtain comprehensive fusion information;

[0145] Perform speech synthesis on the comprehensive fusion information to obtain speech synthesis information.

[0146] The speech synthesis information generation device provided by the embodiment of the present invention can further accurately and efficiently recognize the speech information that a language-impaired customer wants to express by performing speech synthesis on the comprehensive fusion information, improving the business handling efficiency.

[0147] Further, the synthesis unit 403 is specifically configured to:

[0148] Perform voice synthesis on the comprehensive fusion information based on a preset voice synthesis model to obtain voice synthesis information;

[0149] Wherein, the preset voice synthesis model is a multi-modal model that can perform voice synthesis on video content and language content respectively.

[0150] The voice synthesis information generation device provided by the embodiments of the present invention can improve the accuracy and processing efficiency of voice synthesis by performing voice synthesis on comprehensive fusion information through a model.

[0151] Further, the voice synthesis information generation device is further configured to:

[0152] Output the voice synthesis information through a speaker.

[0153] The voice synthesis information generation device provided by the embodiments of the present invention can enable bank staff and customers to hear voice information in a timely manner, further facilitating business handling.

[0154] The embodiments of the voice synthesis information generation device provided by the present invention can specifically be used to execute the processing flows of the above method embodiments. Its functions will not be elaborated here and can refer to the detailed descriptions of the above method embodiments.

[0155] Figure 5 It is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention. As Figure 5 shown, the electronic device includes: a processor 501, a memory 502, and a bus 503;

[0156] Wherein, the processor 501 and the memory 502 communicate with each other through the bus 503;

[0157] The processor 501 is used to call program instructions in the memory 502 to execute the methods provided by the above method embodiments, for example, including:

[0158] Obtain a target video, and extract human joint point features and lip language features according to the target video; the target video includes real-time images of customers with speech impairments;

[0159] Fuse the human joint point features and the lip language features to obtain fused features, and perform text recognition on the fused features to obtain text vocabulary information;

[0160] Perform voice synthesis on the text vocabulary information to obtain voice synthesis information.

[0161] This embodiment discloses a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above method embodiments, for example, including:

[0162] Obtain a target video, and extract human joint point features and lip language features according to the target video; the target video includes real-time images of customers with language barriers;

[0163] Fuse the human joint point features and the lip language features to obtain a fused feature, and perform text recognition on the fused feature to obtain text vocabulary information;

[0164] Perform speech synthesis on the text vocabulary information to obtain speech synthesis information.

[0165] This embodiment provides a computer-readable storage medium, which stores a computer program that enables the computer to execute the methods provided in the above method embodiments, for example, including:

[0166] Obtain a target video, and extract human joint point features and lip language features according to the target video; the target video includes real-time images of customers with language barriers;

[0167] Fuse the human joint point features and the lip language features to obtain a fused feature, and perform text recognition on the fused feature to obtain text vocabulary information;

[0168] Perform speech synthesis on the text vocabulary information to obtain speech synthesis information.

[0169] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0170] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0171] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0172] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0173] In the description of this specification, the description with reference to terms such as "one embodiment", "a specific embodiment", "some embodiments", "for example", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0174] The above specific embodiments have further elaborated on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for generating speech synthesis information, characterized in that, Including: Obtain a target video, and extract human joint point features and lip language features according to the target video; The target video contains real-time images of customers with speech disorders; Fuse the human joint point features and the lip language features to obtain a fused feature, and perform text recognition on the fused feature to obtain text vocabulary information; Perform speech synthesis on the text vocabulary information to obtain speech synthesis information; Wherein, the human joint point features include finger joint features; correspondingly, the method for generating speech synthesis information further includes: if a target object pointed to by a finger is recognized based on the finger joint features, extract the target object features; fuse the human joint point features, the lip language features and the target object features to obtain the fused feature; Wherein, the performing text recognition on the fused feature to obtain text vocabulary information includes: performing text recognition on the fused feature based on a preset feature recognition model to obtain text vocabulary information; wherein, the preset feature recognition model is obtained by training a neural network according to feature recognition sample data; Wherein, before the step of performing text recognition on the fused feature based on the preset feature recognition model, the method for generating speech synthesis information further includes: performing sentence segmentation detection on the target video to obtain sentence segmentation nodes, and segmenting the target video with the sentence segmentation nodes to obtain each target video segment; performing text recognition on the fused features in each target video segment respectively based on the preset feature recognition model.

2. The method for generating speech synthesis information according to claim 1, wherein The method for generating speech synthesis information further includes: Fuse the text vocabulary information and the target video to obtain comprehensive fusion information; Perform speech synthesis on the comprehensive fusion information to obtain speech synthesis information.

3. The method for generating speech synthesis information according to claim 2, wherein The performing speech synthesis on the comprehensive fusion information to obtain speech synthesis information includes: performing speech synthesis on the comprehensive fusion information based on a preset speech synthesis model to obtain speech synthesis information; Wherein, the preset speech synthesis model is a multi-modal model that can perform speech synthesis on video content and language content respectively.

4. The method for generating speech synthesis information according to claim 1, wherein After the step of obtaining speech synthesis information, the method for generating speech synthesis information further includes: Output the speech synthesis information through a speaker.

5. A voice synthesis information generation device, characterized in that, Including: An acquisition unit, configured to acquire a target video, and extract human joint point features and lip language features according to the target video; The target video contains real-time images of customers with speech disorders; A fusion unit, configured to fuse the human joint point features and the lip language features to obtain a fused feature, and perform text recognition on the fused feature to obtain text vocabulary information; A synthesis unit, configured to perform speech synthesis on the text vocabulary information to obtain speech synthesis information; Wherein, the human joint point features include finger joint features; correspondingly, the method for generating speech synthesis information further includes: if a target object pointed to by a finger is recognized based on the finger joint features, extract the target object features; fuse the human joint point features, the lip language features and the target object features to obtain the fused feature; Among them, performing text recognition on the fusion feature to obtain text vocabulary information includes: performing text recognition on the fusion feature based on a preset feature recognition model to obtain text vocabulary information; wherein, the preset feature recognition model is obtained by training a neural network according to feature recognition sample data. Among them, before the step of performing text recognition on the fusion feature based on the preset feature recognition model, the speech synthesis information generation method further includes: performing sentence segmentation detection on the target video to obtain sentence segmentation nodes, and segmenting the target video with the sentence segmentation nodes to obtain each target video segment; performing text recognition on the fusion feature in each target video segment based on the preset feature recognition model.

6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Children reading and interacting method and system based on an intelligent robot

    CN109522835A

  • Semantic recognition method based on human body action analysis and related device

    CN114155606A