Intelligent voice AI soothing method
By simulating the voice and language of the deceased and using intelligent voice AI methods, the problem of lack of continuous comfort for the deceased's family members is solved, and the emotional relief and comfort effects of the deceased's family members are achieved.
Patent Information
- Application Number
- CN202111509354.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The family of the deceased lacks continuous comfort after the death of their loved one, which makes it difficult to relieve grief and makes it impossible for close people to provide long-term comfort.
By obtaining the deceased's recordings and call recordings, extracting the timbre and speaking style, combining machine learning to simulate the deceased's voice and language, and using intelligent voice AI to simulate conversations between the deceased and their family members, the system creates the illusion that the deceased is still alive.
When the family of the deceased talks to AI, it simulates the voice and language of the deceased, relieves sadness, and provides the illusion that the deceased is still alive, achieving a soothing effect.
Smart Images

Figure CN114168713B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to an intelligent voice AI pacification method. BACKGROUND
[0002] Death is a common emotion of human beings, and the death of a relative will cause the family members to feel sad, and they often need to be comforted by other people, so that the sad emotions will improve. However, once the comforting behavior is stopped, the emotions of the bereaved family members will again fall into sadness. There is a phenomenon in the comforting work for the bereaved family members, that is, the closer the relationship, the better the comforting effect. However, the close people cannot have enough time to comfort the bereaved family members, so the sad emotions of the bereaved family members usually have to be digested by themselves, so as to completely eliminate. SUMMARY
[0003] The purpose of the present application is to overcome the problems existing in the prior art, and to provide an intelligent voice AI pacification method, which simulates the speech of the deceased by using the speech manner and voice of the deceased in life, provides the bereaved family members with the illusion that the deceased is still alive, and thus relieves the sad emotions of the bereaved family members.
[0004] To this end, the present application provides an intelligent voice AI pacification method, comprising the following steps:
[0005] Step 1: obtaining the original recording of the deceased in the deceased's mobile phone and the conversation recording of the deceased;
[0006] Step 2: extracting the tone of the original recording and the tone of the conversation recording to obtain the pacification tone of the deceased;
[0007] Step 3: obtaining the conversation between the deceased and the family members before death, training the conversation by machine learning, the conversation including the speech of the deceased and the speech of the family members, and obtaining the speaking manner of the deceased;
[0008] Step 4: receiving the speech of the family members, obtaining the speech of the deceased according to the speaking manner, and playing the obtained speech of the deceased using the pacification tone.
[0009] Further, in step 3, the following steps are included:
[0010] Step 3-1: obtaining the conversation between the deceased and the family members before death, and disassembling the conversation into a plurality of groups of short speeches, each group of short speeches including an adjacent speech of the deceased and a speech of the family members, the speech of the family members being located before the speech of the deceased;
[0011] Step 3-2: using a text extraction technique to extract a keyword set of each of the speeches of the deceased and a keyword set of each of the speeches of the family members, respectively;
[0012] Step 3-3: Establishing a speech model, which is a supervised learning model, taking the keyword set of the deceased speech as output and the keyword set of the family speech as input, and training the speech model by traversing several groups of short speeches;
[0013] Step 3-4: Taking the trained speech model as the speaking manner of the deceased.
[0014] Further, in step 3-3, the following steps are included:
[0015] Step 3-3-1: Converting the keyword set of the deceased speech and the keyword set of the family speech into deceased array and family array respectively, each keyword corresponding to a numerical value;
[0016] Step 3-3-2: Training the speech model by taking the family array as input and the deceased array as output;
[0017] Step 3-3-3: Traversing several groups of short speeches using step 3-3-1 and step 3-3-2 in turn to obtain the trained speech model.
[0018] Further, in step 2, the following steps are included:
[0019] Step 2-1: Weighting the timbre of the original recording and the timbre of the call recording to obtain a machine timbre;
[0020] Step 2-2: Obtaining the facial image of the deceased through image acquisition, obtaining the facial features of the family using image processing technology, and obtaining the simulated timbre according to the facial features of the family;
[0021] Step 2-3: Weighting the machine timbre and the simulated timbre to obtain the soothing timbre.
[0022] Further, in step 2-2, the facial features are facial contours, and when obtaining the simulated timbre according to the facial features of the family, the following steps are included:
[0023] Step 2-2-1: Establishing a two-dimensional coordinate system, importing the facial contours into the two-dimensional coordinate system, and obtaining the coordinates of the facial contours;
[0024] Step 2-2-2: Dividing the facial contours into line segments, obtaining the coordinate set of each line segment and the position of each line segment, obtaining the function of each line segment according to the coordinate set, and corresponding each line segment position with a function;
[0025] Step 2-2-3: According to the function corresponding to each position, find the corresponding factor in the timbre database as the simulated timbre output; the timbre database is used to store the position range of each timbre corresponding line segment and the function corresponding to the position range.
[0026] Further, the position of the line segment is obtained through the coordinates corresponding to the line segment, including the following steps:
[0027] Step 1: Obtain each coordinate of the line segment, and decompose each coordinate into horizontal and vertical coordinates;
[0028] Step 2: Calculate the average of the horizontal coordinates of each coordinate in the line segment, and calculate the average of the vertical coordinates of each coordinate in the line segment;
[0029] Step 3: Combine the average of the horizontal coordinates of each coordinate and the average of the vertical coordinates of each coordinate to obtain the position of the line segment.
[0030] Further, when the face contour is divided into each line segment, the following steps are included:
[0031] Step (1): Randomly retrieve two adjacent coordinates, and use MATLAB to obtain the function corresponding to the adjacent coordinates;
[0032] Step (2): Obtain the critical coordinates of the adjacent coordinates, and input the critical coordinates into the function corresponding to the adjacent coordinates to obtain the critical error;
[0033] Step (3): When the critical error is within the set range, sequentially input the critical coordinates in the sequential direction to perform step (2), cache all coordinates, until the critical error exceeds the set range, and output the coordinate set;
[0034] Step (4): When the critical error exceeds the set range, obtain the critical coordinates and the adjacent coordinates in the sequential direction, and perform step (1);
[0035] Step (5): Obtain the corresponding line segment according to the obtained coordinate set, and output.
[0036] Further, in step 2, the present application further provides a timbre adjusting module, which adjusts the soothing timbre by accepting the frequency and waveform adjusted by the user, and updates the soothing timbre.
[0037] The intelligent voice AI soothing method provided by the present application has the following beneficial effects:
[0038] The application simulates the tone of the deceased by collecting the collected recording of the deceased and the recording of the deceased's previous conversation, and trains the language model of the deceased by combining the speaking tone of the deceased, so that the deceased's family uses the deceased's voice to simulate the deceased's speaking manner and replies to the deceased's family in language when the deceased's family talks with the deceased;
[0039] The application extracts keywords from a large number of conversations of the deceased before death, and trains all the extracted keywords to the language model, so that the speech model completely corresponds to the deceased, so that the deceased's family uses the deceased's voice to simulate the deceased's speaking manner and replies to the deceased's family in language when the deceased's family talks with the deceased;
[0040] The application combines the recording and conversation voice of the mobile phone to obtain the machine voice when simulating the voice tone of the deceased, and trains the simulated voice of the deceased by combining the facial features of the deceased, and obtains the voice tone of the deceased by weighting, and finally adjusts the voice of the deceased by the family, and adjusts the voice of the deceased by inputting the degree of the voice to be strengthened and recording the example of the strengthened voice. BRIEF DESCRIPTION OF DRAWINGS
[0041] Figure 1 The present application is a whole process schematic diagram;
[0042] Figure 2 The present application is a flowchart for obtaining the position of the line segment;
[0043] Figure 3 The present application is a flowchart for dividing the facial contour into line segments;
[0044] Figure 4 The present application is a schematic diagram for dividing the facial contour into line segments. DETAILED DESCRIPTION
[0045] The present application is described in detail below in combination with the drawings, but it should be understood that the protection scope of the present application is not limited by the specific embodiments.
[0046] In the present application, the component model and structure without explicit indication are the prior art known by those skilled in the art, and those skilled in the art can set according to the actual situation, and the specific limitation is not made in the embodiment of the present application.
[0047] Specifically, as shown in the figure, Figures 1-4 The present application provides an intelligent voice AI soothing method, which comprises the following steps:
[0048] Step 1: obtaining the original recording of the deceased in the deceased's mobile phone and the conversation recording of the deceased;
[0049] Step 2: extract the timbre of the original recording and the timbre of the call recording, and synthesize the soothing timbre of the deceased;
[0050] Step 3: obtain the conversation between the deceased and the family members before the death, train the conversation by machine learning, the conversation includes the deceased's speech and the family members' speech, and obtain the speaking manner of the deceased;
[0051] Step 4: receive the family members' speech, get the deceased's speech according to the speaking manner, and play the obtained deceased's speech using the soothing timbre.
[0052] The present application proposes the above steps 1-4, which are sequentially performed in order, in step 1, the deceased's mobile phone is effectively used to retrieve the original recording in the deceased's mobile phone and the deceased's call recording, the original recording in the mobile phone refers to the recording audio saved in the mobile phone, and the call recording refers to the recording of the conversation when the deceased talks with others, wherein the call recording can be obtained from a communication company. In step 2, the two are combined to obtain a soothing timbre that can simulate the sound of the deceased, which can be considered as the timbre of the deceased's speech, so that the family members have the feeling of the deceased speaking when they listen, and the deceased's speech is restored from the perspective of sound, in step 3, the speaking manner of the deceased is obtained, the present application uses machine learning to restore the speaking manner (also the speaking habit) of the deceased according to the past conversation of the deceased, so that the family members can remember the deceased from the auditory and sensory aspects, thereby achieving the effect of comforting the family members, and step 4 is the playing process, the above obtained characteristics are released and executed by playing.
[0053] Therefore, the present application simulates the timbre of the deceased by collecting the recording in the collection used by the deceased and the recording of the call before the death of the deceased, trains the language model of the deceased by combining the speaking tone of the deceased before, so that the family members of the deceased use the voice of the deceased to simulate the speaking manner of the deceased when they have a conversation with the deceased, and reply to the family members of the deceased in language.
[0054] In addition, the present application can also set the carrier as a robot, the built-in sound in the robot can execute the content in step 4, and the robot can also have a face, which is set on the head of the robot in the form of an arc-shaped display to display the appearance of the deceased, and furthermore, a one-to-one restored model can be made according to the appearance of the deceased, and the present application is embedded in the model, thereby achieving the effect of comforting the family members of the deceased.
[0055] In the embodiment of the present application, in the above step 3, when specifically implemented, the following steps are included:
[0056] Step 3-1: obtaining the conversation between the deceased and the family members before the deceased passed away, and disassembling the conversation into a plurality of groups of short conversations, each group of short conversations including an adjacent deceased speech and a family member speech, the family member speech being located before the deceased speech;
[0057] Step 3-2: respectively extracting a keyword set of each of the deceased speeches and a keyword set of each of the family member speeches using a text extraction technique;
[0058] Step 3-3: establishing a speech model, the speech model being a supervised learning model, the keyword set of the deceased speech being taken as an output, the keyword set of the family member speech being taken as an input, and the speech model being trained by traversing a plurality of groups of the short conversations;
[0059] Step 3-4: taking the trained speech model as the speaking manner of the deceased.
[0060] The present application uses a learning model in artificial intelligence, trains a brand-new learning model through the conversation between the deceased and the family members before the deceased passed away, so that the learning model and the deceased are one-to-one corresponding, and the learning model can be used to describe and express the speaking manner of the deceased. In the present application, the learning model used is a supervised learning model.
[0061] In the present application, there are different personnel in the family, and the conversation between the deceased and each family member is extracted and trained respectively. When the trained learning model is used subsequently, different training models are used according to different family members.
[0062] In the present application, in step 3-2, each speech includes a plurality of keywords, which can be obtained using a text extraction technique. When arranging, the keywords are arranged in order according to the order of the keywords in the speech.
[0063] As an optimization of the above technical solution, in step 3-3, the following steps are included:
[0064] Step 3-3-1: respectively converting the keyword set of the deceased speech and the keyword set of the family member speech into a deceased array and a family member array, each keyword corresponding to a numerical value;
[0065] Step 3-3-2: taking the family member array as an input and the deceased array as an output to train the speech model;
[0066] Step 3-3-3: sequentially using step 3-3-1 and step 3-3-2 to traverse a plurality of groups of the short conversations to obtain the trained speech model.
[0067] The above is the process of converting keywords and values, steps 3-3-1 to 3-3-3 are performed in turn, and each keyword corresponds to a value which is found through a keyword database, which is used to store keywords and corresponding values, so that the deceased's speech and the family's speech can be represented using arrays respectively, and when training the learning model, the model can be trained in the form of numbers, which effectively increases the intensity of training. When playing later, the above conversion process can be reversed.
[0068] In an embodiment of the present application, when restoring the timbre of the deceased's speaking voice, the present application also adds the element of the deceased's facial features, which is to make the restoration of the deceased's timbre more realistic, that is, step 2 above can be broken down as follows:
[0069] Step 2-1: weighting the timbre of the original recording and the timbre of the conversation recording to obtain a machine timbre;
[0070] Step 2-2: obtaining the deceased's facial image through image acquisition, using image processing technology to obtain the family's facial features, and obtaining a simulated timbre according to the family's facial features;
[0071] Step 2-3: weighting the machine timbre and the simulated timbre to obtain the soothing timbre.
[0072] The above steps 2-1 to 2-3 are performed in turn, in existing technology, people with different facial features have different timbres, for example, the timbre of some people with wide faces is generally deeper, therefore, researchers at the Massachusetts Institute of Technology have invented a technology that can depict a face from sound alone, therefore, the reverse is also true. By adding the deceased's facial features, the timbre of the deceased can be restored from multiple perspectives in terms of air propagation. At the same time, since the original recording and the conversation recording are both electronic recordings, they are compressed for storage and transmission convenience when recording, which causes distortion of the sound, therefore, by adding the deceased's facial features, the sound of the deceased can be restored to a high degree, achieving a realistic effect.
[0073] As an optimization of the above technical solution, in step 2-2, the facial features are facial contours, which can be obtained from the deceased's facial image through image processing, for example, current facial recognition is completed by comparing the facial contours extracted from the facial image, based on the above method, when obtaining a simulated timbre according to the family's facial features, the present application includes the following steps:
[0074] Step 2-2-1: a two-dimensional coordinate system is established, the face contour is introduced into the two-dimensional coordinate system, and each coordinate of the face contour is obtained;
[0075] Step 2-2-2: the face contour is divided into each line segment, and the coordinate set of each line segment and the position of each line segment are obtained, the function of each line segment is obtained according to the coordinate set of each line segment, and the position of each line segment is corresponded to the function one by one;
[0076] Step 2-2-3: according to the function corresponding to each position, the corresponding factor in the tone database is searched as the simulated tone output; the tone database is used to store the position range of each tone corresponding to the line segment and the function corresponding to the position range.
[0077] The above steps 2-2-1 to 2-2-3 are sequentially performed in order, the functions of the face contour are obtained by using the face contour coordinates, and the corresponding tone in the tone database is searched as the data function according to the position of each function. In the present application, step 2-2-1 is the process of obtaining coordinates, step 2-2-2 is to obtain each function according to the coordinates, and to determine the position of the coordinates according to the obtained function and coordinates, and step 2-2-3 is the process of searching and calling data, so as to obtain the simulated tone.
[0078] The present application uses functions and positions of functions to represent the face contour, and combines the functions with the position relationship. Since the function is an approximate function, the approximate function is that the coordinates are located on both sides of the function with uniform coordinates. Therefore, the obtained face contour is uniform, that is, different people may have the same face contour, and therefore the obtained tone is relatively accurate.
[0079] As an optimization of the above technical solution, the position of the line segment is obtained by the coordinates corresponding to the line segment, including the following steps:
[0080] Step one: obtain each coordinate of the line segment, and decompose each coordinate into horizontal coordinate and vertical coordinate;
[0081] Step two: calculate the average of the horizontal coordinates of each coordinate in the line segment, and calculate the average of the vertical coordinates of each coordinate in the line segment;
[0082] Step three: combine the average of the horizontal coordinates of each coordinate and the average of the vertical coordinates of each coordinate to obtain the position of the line segment.
[0083] The above steps are to obtain the position of the line segment, that is, to obtain the average of the coordinates corresponding to each line segment. In this way, the position of the line segment can be obtained in the curved line segment or the ratio line segment, and the position of the line segment is single.
[0084] As the optimization of the above technical solution, when the face profile is divided into each line segment, the following steps are included:
[0085] Step (1): randomly call two adjacent coordinates, and use MATLAB to obtain the function corresponding to the adjacent coordinates;
[0086] Step (2): obtain the critical coordinates of the adjacent coordinates, and bring the critical coordinates into the function corresponding to the adjacent coordinates to obtain the critical error;
[0087] Step (3): when the critical error is within the set range, sequentially bring the critical coordinates in the sequential direction to execute step (2), cache all the coordinates, until the critical error exceeds the set range, and output the coordinate set;
[0088] Step (4): when the critical error exceeds the set range, obtain the critical coordinates and the adjacent coordinates in the sequential direction, and execute step (1);
[0089] Step (5): according to the obtained coordinate set, the corresponding line segment is obtained and output.
[0090] In the above steps (1) to (5), the process of obtaining each line segment and the corresponding function of the face profile is used. In the execution, the approximate function is used, that is, the coordinates are uniformly located on the left and right sides of the function, combined with Figure 4 , the circles with labels 1-4 are coordinates, in step (1), the coordinates with labels 1 and 2 are selected to obtain the function as the dashed line in the figure, in step (2), the critical coordinates are the coordinates with labels 1 and 2, that is, the coordinate with label 3. Since the coordinate with label 3 is a critical coordinate, it is brought into the function of the dashed line, with the horizontal coordinate as the independent variable and the vertical coordinate as the dependent variable. The difference between the result obtained by bringing the horizontal coordinate into the function and the vertical coordinate is the critical error. Obviously, the coordinate with label 3 belongs to the function of the dashed line in the figure, so the coordinates with labels 1, 2 and 3 and the function of the dashed line are saved correspondingly. Similarly, the sequential direction is the direction of the arrow in the figure. This direction can be counterclockwise or clockwise. The coordinate with label 4 is operated in step (2) to obtain a critical error greater than the set range. Therefore, the coordinate with label 4 does not belong to the function of the dashed line in the figure, so step (4) is executed to match the function of the line segment with the dashed line in the figure for the coordinate with label 4. Step (5) is the output process of the function and the corresponding coordinates. According to the sequential direction, each coordinate in the face profile is sequentially traversed to obtain the result.
[0091] In the embodiment of the present application, in step 2, the present application is also provided with a timbre adjusting module, which adjusts the soothing timbre by accepting the frequency and waveform adjusted by the user, and updates the soothing timbre. The present application obtains the sound closest to the voice of the deceased by means of self-adjustment of the family, thereby producing the greatest degree of comfort for the family.
[0092] The above disclosed are only several specific embodiments of the present application, but the embodiments of the present application are not limited thereto, and any changes that can be thought of by those skilled in the art shall fall within the protection scope of the present application.
Claims
1. An intelligent voice AI soothing method, characterized in that, The method comprises the following steps: Step 1: obtaining the original recording of the deceased in the deceased's mobile phone and the conversation recording of the deceased; Step 2: extracting the tone of the original recording and the tone of the conversation recording to obtain the soothing tone of the deceased; Step 3: obtaining the conversation between the deceased and the family members before the death of the deceased, training the conversation by machine learning, the conversation including the deceased's speech and the family member's speech, and obtaining the speaking manner of the deceased; Step 4: receiving the family member's speech, obtaining the deceased's speech according to the speaking manner, and playing the obtained deceased's speech using the soothing tone; In step 2, the method comprises the following steps: Step 2-1: weighting the tone of the original recording and the tone of the conversation recording to obtain a machine tone; Step 2-2: obtaining the facial image of the deceased by image, obtaining the facial features of the family members using image processing technology, and obtaining the simulated tone according to the facial features of the family members; Step 2-3: weighting the machine tone and the simulated tone to obtain the soothing tone; In step 2-2, the facial features are facial contours, and when obtaining the simulated tone according to the facial features of the family members, the method comprises the following steps: Step 2-2-1: establishing a two-dimensional coordinate system, importing the facial contours into the two-dimensional coordinate system, and obtaining the coordinates of the facial contours; Step 2-2-2: dividing the facial contours into line segments, obtaining the coordinate set of each line segment and the position of each line segment, obtaining the function of each line segment according to the coordinate set, and corresponding each line segment position with a function; Step 2-2-3: according to the function corresponding to each position, searching for the corresponding factor in the tone database as the simulated tone output; the tone database is used to store the position range of each tone corresponding to the line segment and the function corresponding to the position range.
2. The intelligent voice AI soothing method of claim 1, wherein, In step 3, the method comprises the following steps: Step 3-1: obtaining the conversation between the deceased and the family members before the death of the deceased, and dividing the conversation into a plurality of groups of short speeches, each group of short speeches including an adjacent deceased speech and a family member speech, and the family member speech being located before the deceased speech; Step 3-2: using a text extraction technology to extract a keyword set of each of the deceased speeches and a keyword set of each of the family member speeches; Step 3-3: establishing a speech model, the speech model being a supervised learning model, the keyword set of the deceased speeches being taken as an output, the keyword set of the family member speeches being taken as an input, and the speech model being trained by traversing a plurality of groups of the short speeches; Step 3-4: taking the trained speech model as the speaking manner of the deceased.
3. The intelligent voice AI soothing method of claim 2, wherein, In step 3-3, the method comprises the following steps: Step 3-3-1: respectively converting the keyword set of the deceased speeches and the keyword set of the family member speeches into a deceased array and a family member array, each keyword corresponding to a numerical value; Step 3-3-2: taking the family member array as an input and the deceased array as an output to train the speech model; Step 3-3-3: traversing several groups of the short words using step 3-3-1 and step 3-3-2 in sequence to obtain the trained speech model.
4. The intelligent voice AI soothing method of claim 1, wherein, The position of the line segment is obtained by the coordinates corresponding to the line segment, comprising the following steps: Step 1: obtaining each coordinate of the line segment, and decomposing each coordinate into an abscissa and an ordinate; Step 2: calculating the average of the abscissa of each coordinate in the line segment, and calculating the average of the ordinate of each coordinate in the line segment; Step 3: combining the average of the abscissa of each coordinate and the average of the ordinate of each coordinate to obtain the position of the line segment.
5. The intelligent voice AI soothing method of claim 1, wherein, When the face contour is divided into line segments, the following steps are included: Step (1): randomly calling two adjacent coordinates, and using MATLAB to obtain the function corresponding to the adjacent coordinates; Step (2): obtaining the critical coordinates of the adjacent coordinates, and bringing the critical coordinates into the function corresponding to the adjacent coordinates to obtain the critical error; Step (3): when the critical error is within the set range, sequentially bring in the critical coordinates in the sequential direction to execute step (2), cache all the coordinates, until the critical error exceeds the set range, and output the coordinate set; Step (4): when the critical error exceeds the set range, obtaining the critical coordinates and the adjacent coordinates in the sequential direction, and executing step (1); Step (5): obtaining the corresponding line segment according to the obtained coordinate set, and outputting.
6. The intelligent voice AI soothing method of claim 1, wherein, In step 2, there is also a tone adjustment module, which adjusts the soothing tone by accepting the frequency and waveform adjusted by the user, and updates the soothing tone.
Citation Information
Patent Citations
Cinerary casket for carrying out man-machine conversation based on simulating sound of departed person by artificial intelligence (AI)
CN111317642A
Dialogue implementation method and system oriented to health care
CN113488057A
Apparatus for realizing coversation with died person
KR1020210015977A