A method, system, device and medium for determining a degree of relationship
Patent Information
- Application Number
- CN202410033877.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2044-01-09
AI Technical Summary
这样辅助方法的缺失使得人们在社交互动中更加依赖直觉和经验,而这种主观性判断可能会导致误判或产生不必要的矛盾纠纷
[0056]Compared with existing technologies, the beneficial effects of this invention are as follows: by initializing and calibrating the user's voice information, it can be personalized to adapt to the user's voice characteristics and habits, help people identify and analyze the closeness and distance features in the voice information, and calculate the degree of closeness and distance between the user and the target object based on the comparison between the frequency of closeness and distance features and the user's pronunciation habit model. Thus, it helps people understand and judge the degree of closeness and distance of the relationship through voice information, assists people in understanding and judging the degree of closeness and distance of the relationship more accurately in the social process, and enables people to make more appropriate communication responses.
Smart Images

Figure CN117935856B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of information technology, and in particular relates to a method, system, device and medium for determining the degree of closeness of a relationship. Background Technology
[0002] The perception and judgment of closeness in interpersonal relationships are crucial to the sequence of social interactions. However, biases in the perception and judgment of closeness between communicators can lead to adverse effects. Currently, people rely primarily on subjective, habitual perceptions when judging closeness, a dependence easily influenced by individuals with weaker expressive abilities or peculiar verbal habits, thus increasing the likelihood of misunderstandings. There is currently no reliable technological method to assist people in judging the degree of closeness. This lack of auxiliary methods forces people to rely more on intuition and experience in social interactions, and such subjective judgments may lead to misjudgments or unnecessary conflicts. Summary of the Invention
[0003] The purpose of this invention is to provide a method, system, device, and medium for judging the closeness of relationships. By using sound information, this invention helps people understand and judge the closeness of relationships, assists people in more accurately understanding and judging the closeness of relationships in social interactions, and enables people to make more appropriate communication responses.
[0004] This invention is achieved through the following technical solution:
[0005] A method for determining the closeness of a relationship includes the following steps:
[0006] Acquire the user's initial voice information when communicating with the target object, and filter the acquired initial voice information;
[0007] The first sound information is input into the constructed sound affinity feature recognition model to obtain affinity feature data in the first sound information. Affinity feature data includes the number of close feature pronunciations and the number of distant feature pronunciations.
[0008] Based on the affinity feature data, the occurrence frequency of affinity feature pronunciation and the occurrence frequency of alienation feature pronunciation in the first sound information are calculated;
[0009] Intimacy and alienation are calculated according to formulas (1) and (2) respectively;
[0010] Yint = P(Xint) * 100 (1);
[0011] Yest = P(Xest) * 100 (2);
[0012] In the formula, Yint represents intimacy, P(Xint) is the cumulative probability that the frequency of occurrence of intimacy features is less than or equal to the value of Xint in the constructed user pronunciation habit model, and Yet represents alienation, P(Xest) is the cumulative probability that the frequency of occurrence of alienation features is less than or equal to the value of Xest in the constructed user pronunciation habit model.
[0013] Based on the calculated intimacy and alienation, the degree of relationship between the user and the target object is determined.
[0014] Furthermore, the step of filtering the acquired first audio information includes:
[0015] The first sound information is noise-reduced to remove interference from surrounding noise;
[0016] Convert the first audio information into text information and the timestamp of the text information;
[0017] Based on a pre-defined filter word list, the text information is filtered, and the timestamps corresponding to the filtered words in the text information are recorded to obtain the target timestamp.
[0018] Extract the time segment corresponding to the target timestamp from the first audio information to obtain the new first audio information.
[0019] Furthermore, the steps for obtaining the user's initial voice information during communication with the target include:
[0020] Collect multiple second voice information from natural dialogues between users and target objects;
[0021] Multiple second audio messages are spliced and integrated in the order of collection time to obtain the first audio message, and the duration of the first audio message is made to a set duration.
[0022] Furthermore, the construction process of the voice affinity feature recognition model is as follows:
[0023] Collect third-voice information from multiple different objects in natural dialogue to obtain multiple third-voice information;
[0024] For each third audio information, noise reduction and segmentation are performed to obtain several audio information segments;
[0025] Obtain the characteristic pronunciation of each character in each audio information segment. The characteristic pronunciation includes close characteristic pronunciation, distant characteristic pronunciation or no characteristic pronunciation. Among them, close characteristic pronunciation is nasal, entering tone, closed mouth, labiodental or blurred sound, and distant characteristic pronunciation is open mouth, guttural sound or retroflex sound.
[0026] The multiple audio information segments were divided into training and validation groups;
[0027] The training group was used to fine-tune the pre-trained voice affinity feature recognition model, and the validation group was used to verify the accuracy and robustness of the pre-trained voice affinity feature recognition model.
[0028] Furthermore, the steps for obtaining the characteristic pronunciation of each character in each audio information segment include:
[0029] Acquire several characteristic pronunciations of each character in each audio information segment, as annotated by multiple evaluators;
[0030] For each audio information segment, the feature annotation consistency coefficient of the audio information segment is calculated, and the audio information segments with a feature annotation consistency coefficient higher than the first preset threshold are marked as audio information segments to be evaluated, thus obtaining multiple target audio information segments to be evaluated.
[0031] The multiple target audio information segments to be evaluated are divided into multiple audio information groups;
[0032] For each group of sound information, obtain the sampling accuracy rate of the expert's annotation of the feature pronunciation of the sound information group, and mark the sound information group with a sampling accuracy rate greater than the second preset threshold as a valid annotation;
[0033] In multiple groups of audio information, the groups with valid annotations are retained, while the remaining groups are deleted, thus obtaining the characteristic pronunciation of each character in each audio information segment.
[0034] Furthermore, the process of constructing the user pronunciation habit model is as follows:
[0035] Based on the characteristic pronunciation of each character in each audio information segment, establish a norm for the sound affinity feature;
[0036] After obtaining the affinity feature data from the first sound information, the affinity feature data is stored.
[0037] Determine whether the number of stored affinity feature data has reached the preset number. If so, generate a user pronunciation habit model based on the stored affinity feature data and voice affinity feature norms.
[0038] Furthermore, before calculating intimacy and alienation according to formulas (1) and (2) respectively, the method also includes:
[0039] Determine whether a user pronunciation habit model has been built;
[0040] If so, the user's pronunciation habit model will be calibrated and updated using the affinity feature data.
[0041] If not, then proceed with the process of building a user pronunciation habit model.
[0042] The present invention also provides a system for determining the degree of closeness in a relationship, comprising:
[0043] The acquisition module is used to acquire the first sound information of the user when communicating with the target object, and to filter the acquired first sound information;
[0044] The input module is used to input the first sound information into the constructed sound affinity feature recognition model to obtain the number of affinity feature pronunciations and aloof feature pronunciations;
[0045] The first calculation module is used to calculate the frequency of occurrence of the proximity feature pronunciation and the frequency of occurrence of the alienation feature pronunciation in the first sound information;
[0046] The second calculation module is used to calculate intimacy and alienation according to formula (1) and formula (2) respectively;
[0047] Yint = P(Xint) * 100 (1);
[0048] Yest = P(Xest) * 100 (2);
[0049] In the formula, Yint represents intimacy, P(Xint) is the cumulative probability that the frequency of occurrence of intimacy features is less than or equal to the value of Xint in the constructed user pronunciation habit model, and Yet represents alienation, P(Xest) is the cumulative probability that the frequency of occurrence of alienation features is less than or equal to the value of Xest in the constructed user pronunciation habit model.
[0050] The confirmation module is used to confirm the degree of relationship between the user and the target object based on calculated intimacy and alienation.
[0051] The present invention also provides an electronic device, comprising:
[0052] processor;
[0053] Memory is used to store executable computer programs;
[0054] The processor executes the computer program to implement the steps of the method for determining the degree of closeness of a relationship.
[0055] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method for determining the degree of closeness of a relationship.
[0056] Compared with existing technologies, the beneficial effects of this invention are as follows: by initializing and calibrating the user's voice information, it can be personalized to adapt to the user's voice characteristics and habits, help people identify and analyze the closeness and distance features in the voice information, and calculate the degree of closeness and distance between the user and the target object based on the comparison between the frequency of closeness and distance features and the user's pronunciation habit model. Thus, it helps people understand and judge the degree of closeness and distance of the relationship through voice information, assists people in understanding and judging the degree of closeness and distance of the relationship more accurately in the social process, and enables people to make more appropriate communication responses. Attached Figure Description
[0057] Figure 1 This is a flowchart illustrating the steps of the method for determining the degree of closeness in a relationship according to the present invention.
[0058] Figure 2 This is a schematic diagram of the modules of the system for determining the degree of closeness of a relationship according to the present invention;
[0059] Figure 3 This is a structural block diagram of an electronic device according to an exemplary embodiment of the present invention. Detailed Implementation
[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0061] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0062] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0063] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0064] In the description of this invention, it should be noted that the terms "upper," "lower," "inner," "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the product of this invention is usually placed when in use. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.
[0065] Please see Figure 1 , Figure 1 This is a flowchart illustrating the steps of the method for determining the degree of closeness in a relationship according to the present invention. A method for determining the degree of closeness in a relationship includes the following steps:
[0066] S1. Obtain the first voice information of the user when communicating with the target object, and filter the obtained first voice information;
[0067] S2. Input the first sound information into the constructed sound affinity feature recognition model to obtain affinity feature data in the first sound information. Affinity feature data includes the number of close feature pronunciations and the number of distant feature pronunciations.
[0068] S3. Based on the affinity feature data, calculate the occurrence frequency of affinity feature pronunciation and the occurrence frequency of alienation feature pronunciation in the first sound information;
[0069] Intimacy and alienation are calculated according to formulas (1) and (2) respectively;
[0070] Yint = P(Xint) * 100 (1);
[0071] Yest = P(Xest) * 100 (2);
[0072] In the formula, Yint represents intimacy, P(Xint) is the cumulative probability that the frequency of occurrence of intimacy features is less than or equal to the value of Xint in the constructed user pronunciation habit model, and Yet represents alienation, P(Xest) is the cumulative probability that the frequency of occurrence of alienation features is less than or equal to the value of Xest in the constructed user pronunciation habit model.
[0073] S4. Based on the calculated intimacy and alienation, confirm the degree of relationship between the user and the target object.
[0074] In step S1 above, the first audio information comes from the voices of the user and the target in a natural conversation, including obtaining it through user-uploaded audio information or by authorizing the system to monitor the voice exchange in real time and capturing the voices of the user and the target in a natural conversation through system monitoring. Then, the obtained first audio information is filtered to remove noise and special words that may interfere with subsequent evaluations.
[0075] Furthermore, in step S1, the step of acquiring the user's first voice information during communication with the target object includes:
[0076] S11. Collect multiple second voice information from the user and the target in natural dialogue;
[0077] S12. Multiple second sound information are spliced and integrated according to the collection time order to obtain first sound information, and the time length of the first sound information is made up to the set time length.
[0078] In steps S11 and S12 above, the acquired first audio information needs to meet a time length requirement. Therefore, conversations between the user and the target object that occur within a similar time frame and use the same language or dialect are collected. The sounds emitted by the user during the conversation are selected and stored to obtain multiple second audio information sets. The similar conversation time frame can be within 5 minutes. Then, the multiple second audio information sets are spliced and integrated according to the collection time order, so that the total time length of the spliced and integrated audio information is greater than or equal to a set time length, preferably 10 minutes.
[0079] Further, in step S1, the step of filtering the acquired first sound information includes:
[0080] S13. Perform noise reduction processing on the first sound information to remove interference from surrounding noise;
[0081] S14. Convert the first audio information into text information and the timestamp of the text information;
[0082] S15. Filter the text information according to the pre-set filter word list, and record the timestamps corresponding to the filtered words in the text information to obtain the target timestamp.
[0083] S16. Remove the time segment corresponding to the target timestamp from the first audio information to obtain new first audio information.
[0084] In steps S13 to S16 above, the first audio information is first subjected to noise reduction processing to filter out noise in the natural dialogue and remove interference from surrounding noise. Then, natural language recognition technology is used to convert the first audio information into text with timestamps, thus obtaining the text information in the first audio information and the time corresponding to each text. Based on a pre-set filtering vocabulary, specific text in the text information is filtered, such as certain verbal tics or specific meaningless words, and the timestamps corresponding to the filtered text are recorded. Then, the recorded timestamps are used to back-filter the audio information in the corresponding time period in the first audio information, that is, to remove irrelevant audio information in the first audio information that may interfere with the evaluation results, avoiding interference with the recognition of affinity features. When the proportion of noise and irrelevant sounds contained in the audio information is lower than a set threshold, the audio information is considered usable for evaluation.
[0085] In step S2 above, a voice affinity feature recognition model is pre-constructed. The input of the voice affinity feature recognition model is voice information, and the output of the voice affinity feature recognition model is affinity feature data in the voice information. The filtered first voice information is then recognized using the constructed voice affinity feature recognition model to extract the number of pronunciations with affinity features and the number of pronunciations with aloof features from the first voice information.
[0086] Furthermore, the construction process of the voice affinity feature recognition model is as follows:
[0087] Step (11): Collect third voice information of multiple different objects in natural dialogue to obtain multiple third voice information;
[0088] Step (12): For each third sound information, perform noise reduction and segmentation on the third sound information to obtain several sound information segments;
[0089] Step (13): Obtain the characteristic pronunciation of each character in each audio information segment. The characteristic pronunciation includes close characteristic pronunciation, distant characteristic pronunciation or no characteristic pronunciation. Among them, close characteristic pronunciation is nasal, entering tone, closed mouth sound, labiodental sound or blurred sound, and distant characteristic pronunciation is open sound, guttural sound or retroflex sound.
[0090] Step (14): Divide the multiple audio information segments into a training group and a validation group;
[0091] Step (15): Use the training group to fine-tune the pre-trained voice affinity feature recognition model, and use the validation group to verify the accuracy and robustness of the pre-trained voice affinity feature recognition model.
[0092] In step (11) above, before constructing the voice affinity feature recognition model, it is necessary to collect multiple sets of third voice information. These sets of third voice information cover the natural dialogues between multiple different objects and people with different relationship types, and are distinguished according to the speech and dialect types spoken by the objects, so as to ensure the accuracy and wide applicability of the model.
[0093] In step (12) above, each third audio message is subjected to noise reduction processing to remove the interference of surrounding noise. Then, each third audio message is split, and the splitting criteria are: retain the integrity of the sentences in the audio message, and the duration of each audio message segment is no more than 30 seconds. The splitting process can be in the form of machine splitting plus manual proofreading.
[0094] In step (13) above, the sound of each character in each audio information segment is evaluated manually, and the characteristic pronunciation of each character is marked, so as to obtain the characteristic pronunciation of each character in each audio information segment.
[0095] Furthermore, in step (13), the step of obtaining the characteristic pronunciation corresponding to the sound of each character in each sound information segment includes:
[0096] Step (131): Obtain several characteristic pronunciations of each character in each audio information segment, as annotated by multiple evaluators;
[0097] Step (132): For each audio information segment, calculate the feature annotation consistency coefficient of the audio information segment, and mark the audio information segments with a feature annotation consistency coefficient higher than the first preset threshold as audio information segments to be evaluated, thereby obtaining multiple target audio information segments to be evaluated;
[0098] Step (133): Divide the multiple target sound information segments to be evaluated into multiple sound information groups;
[0099] Step (134): For each group of sound information, obtain the sampling accuracy of the expert's annotation of the feature pronunciation of the sound information group, and mark the sound information group with a sampling accuracy greater than the second preset threshold as a valid annotation;
[0100] Step (135): In multiple groups of sound information, retain the effective labeled sound information groups and delete the rest of the sound information groups to obtain the characteristic pronunciation label of each character in each sound information segment.
[0101] In steps (131) to (135) above, several trained evaluators evaluate the sound of each character in each audio segment and annotate several characteristic pronunciations corresponding to the sound of each character, thereby obtaining several characteristic pronunciations corresponding to the sound of each character in each audio segment. At least three evaluators are required. Then, the feature annotation consistency coefficient Fleiss'kappa for each audio segment is calculated. The calculation process is existing technology and will not be repeated here. When the feature annotation consistency coefficient Fleiss'kappa of an audio segment is higher than a first preset threshold, it enters the expert evaluation stage. Therefore, audio segments with feature annotation consistency coefficients higher than the first preset threshold are marked as audio segments to be evaluated, where the first preset threshold can be set to 0.8. In the expert evaluation stage, the audio segments to be evaluated are divided into multiple audio groups, each audio group including multiple audio segments. Then, the audio segments to be evaluated within each audio information group are sampled to confirm the correctness of the characteristic pronunciation corresponding to the sound of each character in the audio segment. The sampling ratio is 10%. When the sampling accuracy of a certain audio information group is greater than a second preset threshold, the characteristic pronunciation of the audio segments in that audio information group is considered correct, and therefore the audio information group is marked as a valid annotation. Finally, the validly annotated audio information groups are retained, and the remaining audio information groups are deleted. Based on the retained audio information groups, for each character in each audio segment, the characteristic pronunciation with the most occurrences is taken as the characteristic pronunciation corresponding to the sound of that character, thus obtaining the characteristic pronunciation corresponding to the sound of each character in each audio segment.
[0102] In steps (14) to (15) above, each retained audio segment is divided into two groups. One group is the training group, which serves as training material to fine-tune the pre-trained audio affinity feature recognition model. Specifically, each audio segment in the training group is used as input, and the characteristic pronunciation of each character in each audio segment is used as output to fine-tune the pre-trained audio affinity feature recognition model, resulting in the final audio affinity feature recognition model. The pre-trained audio affinity feature recognition model can use, but is not limited to, wav2vec 2.0. The other group is the validation group, which serves as validation material to verify the accuracy and robustness of the trained audio affinity feature recognition model, ensuring that the audio affinity feature recognition model can stably recognize most audio features.
[0103] In step S2 above, the filtered first sound information is input into the trained sound affinity feature recognition model. The number of affinity feature pronunciations in the first sound information is analyzed. For the sound of each character, if the confidence score exceeds 0.9, the recognition is considered successful. For example, if the sound affinity feature recognition model outputs nasal: 0.7, indistinct sound 0.2, labiodental sound 0.1, the recognition result is considered nasal with a confidence score of 0.7. If it is less than 0.9, the recognition is considered unsuccessful, and the sound of that character is marked as having no pronunciation feature. If the sound affinity feature recognition model outputs nasal: 0.91, indistinct sound 0.05, labiodental sound 0.04, the recognition structure is considered nasal with a confidence score of 0.91. If it exceeds 0.9, the recognition is considered successful, and the sound of that character is marked as nasal. Then, the number of affinity and alienation feature pronunciations in the first sound information is recorded, thereby obtaining the number of affinity and alienation feature pronunciations in the first sound information.
[0104] In step S3 above, the frequency of the familiar characteristic pronunciation, Xint, is calculated as: the number of familiar characteristic pronunciations in the first sound information / the total number of syllables in the first sound information. The number of familiar characteristic pronunciations in the first sound information is the frequency of (nasal sounds + entering sounds + closed sounds + labiodental sounds + indistinct sounds). The frequency of the unfamiliar characteristic pronunciation, Xest, is calculated as: the number of unfamiliar characteristic pronunciations in the first sound information / the total number of syllables in the first sound information. The number of unfamiliar characteristic pronunciations in the first sound information is the frequency of (open sounds + guttural sounds + retroflex sounds).
[0105] In step S4 above, a user pronunciation habit model is pre-constructed. The purpose of this model is to evaluate the user's personalized pronunciation habits in daily conversations, that is, the distribution of the user's pronunciation habits, which is used for the subsequent calculation of the closeness / distance score. Then, the calculated frequency of closeness-feature pronunciation and the frequency of distance-feature pronunciation are compared with the user pronunciation habit model to calculate the closeness / distance rank of the target object in the user's pronunciation habits, and the final closeness / distance score is obtained. That is, the closeness and distance are calculated using formula (1) and formula (2) respectively. The calculation of the cumulative distribution function values P(Xint) and P(Xest) of the frequency of closeness-feature pronunciation and the frequency of distance-feature pronunciation in the constructed user pronunciation habit model is existing technology and will not be elaborated here.
[0106] Furthermore, the process of constructing the user pronunciation habit model is as follows:
[0107] Step (21): Based on the characteristic pronunciation of each character in each audio information segment, establish a norm for the sound affinity feature;
[0108] Step (22): After obtaining the affinity feature data in the first sound information, store the affinity feature data;
[0109] Step (23): Determine whether the number of stored affinity feature data has reached the preset number. If so, generate a user pronunciation habit model based on the stored affinity feature data and voice affinity feature norms.
[0110] In steps (21) to (23) above, based on the characteristic pronunciation of each character in each audio information segment obtained in steps (11), (12) and (13), a user's voice affinity feature norm is established. The voice affinity feature norm is the average usage degree of each characteristic pronunciation in natural dialogue, that is, the mean and standard deviation of each characteristic pronunciation. When the user pronunciation habit model has not yet been generated, the affinity feature data obtained in step S2 is stored first, and the stored affinity feature data is accumulated. The preset number can be selected according to the actual situation. The preset number should not be too small, as too small a number will lead to an inaccurate user pronunciation habit model. It should also not be too large, as too large a number will bring estimation difficulties and interference problems caused by early data. In this embodiment, the lower limit of the preset number is 5 sets, and the upper limit is 50 sets. When the number of stored affinity feature data reaches the preset number, the user's pronunciation habit distribution on various characteristic pronunciations can be estimated by combining the established voice affinity feature norm with initialization processing. That is, the user's pronunciation habit features are initialized to obtain the initial level of the user's pronunciation habits on various characteristic pronunciations, thereby obtaining the user's pronunciation habit model. Specifically, the norm of voice affinity features is used as the prior distribution for Bayesian inference, and the stored affinity feature data is used as the observation condition to obtain a posterior distribution, which is the user pronunciation habit model.
[0111] Furthermore, before step S4, i.e. before calculating intimacy and alienation according to formulas (1) and (2) respectively, the method further includes:
[0112] S4a. Determine whether a user pronunciation habit model has been built;
[0113] S4b. If so, then use the affinity feature data to calibrate the user pronunciation habit model and update the user pronunciation habit model.
[0114] S4c, If not, then proceed with the process of building the user's pronunciation habit model.
[0115] In steps S4a to S4c above, it can be first determined whether a user pronunciation habit model has been generated. This allows for a comparison between the affinity feature data in the obtained first sound information and the user pronunciation habit model to calculate the degree of closeness and distance between the user and the target object. If a user pronunciation habit model has been generated, the affinity feature data in the first sound information obtained in step S2 is used as calibration data to calibrate and update the user pronunciation habit model. As the number of user conversations increases, the amount of conversational sound information available for measurement also increases, making the user pronunciation habit model more accurate.
[0116] In step S5 above, the average score for intimacy and distance is 50 points. Within a certain time period, i.e., the time when the first audio information is recorded, if the user and the target object communicate with a higher degree of intimacy, the higher the intimacy score, the closer the relationship between the user and the target object is likely; if the user and the target object communicate with a higher degree of distance, the higher the distance score, the more distant the relationship between the user and the target object is likely. Specifically, by comparing intimacy and distance, if the intimacy value is higher than the distance value, it indicates that the relationship between the user and the target object is likely close, with the intimacy value representing the degree of closeness; if the distance value is higher than the intimacy value, it indicates that the relationship between the user and the target object is likely distant, with the distance value representing the degree of distance.
[0117] Corresponding to the aforementioned embodiments of the method for determining the degree of closeness of a relationship, the present invention also provides a system for determining the degree of closeness of a relationship, which can be applied to a terminal. For example... Figure 2 As shown, the system includes:
[0118] Acquisition module 1 is used to acquire the first sound information of the user when communicating with the target object, and to filter the acquired first sound information;
[0119] Input module 2 is used to input the first sound information into the constructed sound affinity feature recognition model to obtain the number of affinity feature pronunciations and aloof feature pronunciations;
[0120] The first calculation module 3 is used to calculate the frequency of occurrence of the proximity feature pronunciation and the frequency of occurrence of the alienation feature pronunciation in the first sound information;
[0121] The second calculation module 4 is used to calculate intimacy and alienation according to formula (1) and formula (2) respectively;
[0122] Yint = P(Xint) * 100 (1);
[0123] Yest = P(Xest) * 100 (2);
[0124] In the formula, Yint represents intimacy, P(Xint) is the cumulative probability that the frequency of occurrence of intimacy features is less than or equal to the value of Xint in the constructed user pronunciation habit model, and Yet represents alienation, P(Xest) is the cumulative probability that the frequency of occurrence of alienation features is less than or equal to the value of Xest in the constructed user pronunciation habit model.
[0125] Module 5 is used to confirm the degree of relationship between the user and the target object based on calculated intimacy and alienation.
[0126] Furthermore, module 1 includes:
[0127] The noise reduction submodule is used to perform noise reduction processing on the first sound information to remove interference from surrounding noise;
[0128] The conversion submodule is used to convert the first audio information into text information and the timestamp of the text information;
[0129] The filtering submodule is used to filter text information according to a pre-defined filter vocabulary and record the timestamps corresponding to the filtered text to obtain the target timestamp.
[0130] The truncation submodule is used to truncate the time segment corresponding to the target timestamp in the first audio information to obtain new first audio information.
[0131] Furthermore, module 1 includes:
[0132] The first collection submodule is used to collect multiple second voice information from the user and the target object in natural dialogue;
[0133] The integration submodule is used to splice and integrate multiple second sound information according to the collection time order to obtain the first sound information, and to make the time length of the first sound information reach the set time length.
[0134] Furthermore, input module 2 includes:
[0135] The second collection submodule is used to collect third voice information of multiple different objects in natural dialogue, and obtain multiple third voice information.
[0136] The splitting submodule is used to perform noise reduction and splitting on each third audio information to obtain several audio information segments;
[0137] The acquisition submodule is used to acquire the characteristic pronunciation of each character in each audio information segment. The characteristic pronunciation includes close characteristic pronunciation, distant characteristic pronunciation or no characteristic pronunciation. Among them, close characteristic pronunciation is nasal, entering tone, closed mouth sound, labiodental sound or blurred sound, and distant characteristic pronunciation is open sound, guttural sound or retroflex sound.
[0138] The sub-module is used to divide multiple audio information segments into training and validation groups;
[0139] The training submodule is used to fine-tune the pre-trained voice affinity feature recognition model using the training group, and to verify the accuracy and robustness of the pre-trained voice affinity feature recognition model using the validation group.
[0140] Furthermore, the acquisition sub-modules include:
[0141] The first acquisition unit is used to acquire several characteristic pronunciations of each character in each audio information segment, which are marked by multiple evaluators respectively.
[0142] The calculation unit is used to calculate the feature annotation consistency coefficient of each audio information segment, and mark the audio information segments with feature annotation consistency coefficients higher than a first preset threshold as audio information segments to be evaluated, thereby obtaining multiple target audio information segments to be evaluated.
[0143] The segmentation unit is used to divide multiple target audio information segments to be evaluated into multiple groups of audio information;
[0144] The second acquisition unit is used to acquire the sampling accuracy rate of the expert's feature pronunciation annotation of each group of sound information, and to mark the sound information group with a sampling accuracy rate greater than the second preset threshold as a valid annotation.
[0145] The retention unit is used to retain the validly labeled audio information groups from multiple audio information groups, while deleting the remaining audio information groups, so as to obtain the characteristic pronunciation of each character in each audio information segment.
[0146] Furthermore, the second computing module 4 includes:
[0147] A submodule is created to establish a norm for sound affinity features based on the sound characteristics of each character in each sound information segment.
[0148] The storage submodule is used to store the affinity feature data after the input module 2 obtains the affinity feature data from the first sound information;
[0149] The generation submodule is used to determine whether the number of stored affinity feature data has reached the preset number. If so, it generates a user pronunciation habit model based on the stored affinity feature data and the voice affinity feature norm.
[0150] Furthermore, the system also includes:
[0151] The judgment module is used to determine whether a user pronunciation habit model has been built;
[0152] The update module is used to determine if the module's judgment is correct. If so, the user's pronunciation habit model is calibrated using the affinity feature data, and the user's pronunciation habit model is updated.
[0153] The execution module is used to determine if the module's condition is negative, and then execute the creation, storage, and generation of submodules.
[0154] The implementation process of the functions and roles of each module, submodule and unit in the above system is detailed in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0155] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The system embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units.
[0156] Corresponding to the aforementioned embodiments of the method for determining the degree of closeness of a relationship, the present invention also provides an electronic device, which may include:
[0157] processor;
[0158] Memory is used to store executable computer programs;
[0159] The processor executes the computer program to implement the steps of the aforementioned method for determining the degree of closeness of a relationship.
[0160] The methods and systems for determining the degree of closeness in relationships provided in this invention can all be applied to the aforementioned electronic devices. Taking software implementation as an example, as a logical system, it is formed by the processor of the electronic device loading corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 3 As shown, except Figure 3 In addition to the processor, memory, network interface, and non-volatile memory shown, the electronic device may also include other hardware, such as a camera module; or, depending on the actual function of the electronic device, it may also include other hardware, which will not be described in detail here.
[0161] Corresponding to the aforementioned embodiments of the method for determining the degree of closeness of a relationship, this embodiment of the invention also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the steps of the aforementioned method for determining the degree of closeness of a relationship.
[0162] Embodiments of the present invention may take the form of a computer program product implemented on one or more storage media containing program code (including but not limited to disk storage, CD-ROM, optical storage, etc.). The computer-readable storage medium may include: permanent or non-permanent removable or non-removable media. The information storage function of the computer-readable storage medium can be implemented by any feasible method or technology. The information may be computer-readable instructions, data structures, program models, or other data.
[0163] Additionally, the computer-readable storage medium includes, but is not limited to: phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or other non-transfer media that can be used to store information accessible by a computing device.
[0164] This invention is not limited to the above-described embodiments. If any modifications or variations to this invention do not depart from the spirit and scope of this invention, and if such modifications and variations fall within the scope of the claims and equivalent technologies of this invention, then this invention also intends to include such modifications and variations.
Claims
1. A method of determining a degree of relationship, characterized by, Includes the following steps: Acquire the first voice information of the user when communicating with the target object, and filter the acquired first voice information; The first sound information is input into the constructed sound affinity feature recognition model to obtain affinity feature data in the first sound information. The affinity feature data includes the number of close feature pronunciations and the number of distant feature pronunciations. Based on the affinity feature data, calculate the occurrence frequency of the affinity feature pronunciation and the affinity feature pronunciation in the first sound information; Intimacy and alienation are calculated according to formulas (1) and (2) respectively; Yint = P(Xint) * 100 (1); Yest = P(Xest) * 100 (2); In the formula, Yint represents intimacy, P(Xint) is the cumulative probability that the frequency of occurrence of the intimacy feature pronunciation is less than or equal to the value of Xint in the constructed user pronunciation habit model, and Yet represents alienation, P(Xest) is the cumulative probability that the frequency of occurrence of the alienation feature pronunciation is less than or equal to the value of Xest in the constructed user pronunciation habit model. Based on the calculated intimacy and alienation, the degree of relationship between the user and the target object is determined.
2. The method of determining a degree of relationship according to claim 1, wherein, The step of filtering the acquired first sound information includes: The first sound information is subjected to noise reduction processing to remove interference from surrounding noise; Convert the first audio information into text information and the timestamp of the text information; Based on a pre-defined filter word list, the text information is filtered, and the timestamps corresponding to the filtered words in the text information are recorded to obtain the target timestamp. The time segment corresponding to the target timestamp is removed from the first audio information to obtain new first audio information.
3. The method of determining a degree of relationship according to claim 1, wherein, The step of obtaining the user's first voice information when communicating with the target object includes: Collect multiple second voice information from the user and the target object during natural dialogue; Multiple second sound information items are spliced and integrated according to the collection time sequence to obtain the first sound information, and the time length of the first sound information is made up to a set time length.
4. The method of determining a degree of relationship according to claim 1, wherein, The construction process of the voice affinity feature recognition model is as follows: Collect third voice information from multiple different objects in natural dialogue to obtain multiple sets of said third voice information; For each of the third audio information segments, noise reduction and segmentation are performed to obtain several audio information segments; Obtain the characteristic pronunciation corresponding to the sound of each character in each of the aforementioned sound information segments. The characteristic pronunciation includes close characteristic pronunciation, distant characteristic pronunciation, or no characteristic pronunciation. The close characteristic pronunciation is nasal, entering tone, closed mouth sound, labiodental sound, or indistinct sound. The distant characteristic pronunciation is open sound, guttural sound, or retroflex sound. The multiple audio information segments are divided into a training group and a validation group; The pre-trained voice affinity feature recognition model was fine-tuned using a training group, and the accuracy and robustness of the pre-trained voice affinity feature recognition model were verified using a validation group.
5. The method of determining a degree of relationship according to claim 4, wherein, The step of obtaining the characteristic pronunciation of each character in each audio information segment includes: Acquire several characteristic pronunciations of each character in each segment of the audio information, as annotated by multiple evaluators; For each audio information segment, the feature annotation consistency coefficient of the audio information segment is calculated, and the audio information segments with feature annotation consistency coefficients higher than a first preset threshold are marked as audio information segments to be evaluated, thus obtaining multiple target audio information segments to be evaluated. The multiple target audio information segments to be evaluated are divided into multiple audio information groups; For each group of sound information, the sampling accuracy rate of the expert's annotation of the feature pronunciation of the sound information group is obtained, and the sound information group with a sampling accuracy rate greater than a second preset threshold is marked as a valid annotation; In the multiple groups of sound information, the validly labeled sound information groups are retained, and the remaining sound information groups are deleted, so as to obtain the characteristic pronunciation of each character in each sound information segment.
6. The method of determining a degree of relationship according to claim 4, wherein, The process of constructing the user pronunciation habit model is as follows: Based on the characteristic pronunciation of each character in each audio information segment, establish a norm for the affinity feature of sound; After obtaining the affinity feature data in the first sound information, the affinity feature data is stored. Determine whether the number of stored affinity feature data has reached a preset number. If so, generate the user pronunciation habit model based on the stored affinity feature data and the voice affinity feature norm.
7. The method of determining a degree of relatedness according to claim 6, wherein, Before the steps of calculating intimacy and distance according to formulas (1) and (2) respectively, the method further includes: Determine whether the user's pronunciation habit model has been constructed; If so, the user pronunciation habit model is calibrated and updated using the affinity feature data; If not, then proceed with the process of building the user pronunciation habit model.
8. A system for determining a degree of relationship, the system comprising: include: The acquisition module is used to acquire the first voice information of the user when communicating with the target object, and to filter the acquired first voice information; The input module is used to input the first sound information into the constructed sound affinity feature recognition model to obtain the number of affinity feature pronunciations and aloof feature pronunciations; The first calculation module is used to calculate the frequency of occurrence of the proximity feature pronunciation and the frequency of occurrence of the alienation feature pronunciation in the first sound information; The second calculation module is used to calculate intimacy and alienation according to formula (1) and formula (2) respectively; Yint = P(Xint) * 100 (1); Yest = P(Xest) * 100 (2); In the formula, Yint represents intimacy, P(Xint) is the cumulative probability that the frequency of occurrence of the intimacy feature pronunciation is less than or equal to the value of Xint in the constructed user pronunciation habit model, and Yet represents alienation, P(Xest) is the cumulative probability that the frequency of occurrence of the alienation feature pronunciation is less than or equal to the value of Xest in the constructed user pronunciation habit model. The confirmation module is used to confirm the degree of relationship between the user and the target object based on calculated intimacy and alienation.
9. An electronic device, comprising: include: processor; Memory is used to store executable computer programs; Wherein, when the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
User relationship deduction method based on base station connection information in call bill big data of users
CN107729940A
Intimacy calculation method, intimacy calculation program and intimacy calculation device
JP2013206389A