Key reference analysis method and device in voice-to-text conversion, medium and product
By building a user's interpersonal knowledge graph and multi-dimensional matching, the semantic confusion problem of third-person pronouns in speech-to-text conversion is solved, achieving more accurate information transmission and improved user experience.
Patent Information
- Application Number
- CN202511098423.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-10-03
AI Technical Summary
In the speech-to-text function, third-person pronouns are difficult to distinguish accurately, resulting in semantic confusion and affecting the accuracy of information transmission.
Build a user's interpersonal relationship knowledge graph, split it into multiple relationship sub-graphs, obtain the referents of third-person pronouns through semantic analysis, sort and match them based on context information and confidence, and obtain the correct text sequence of third-person pronouns.
It improves the accuracy of speech-to-text conversion, ensures the clarity of information transmission, and enhances the user experience.
Smart Images

Figure CN120748408A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of natural language processing technology, and in particular to a method, device, medium, and product for analyzing key references in speech-to-text conversion. Background Art
[0002] With the widespread adoption of mobile internet, voice-to-text features have emerged in social applications, providing users with a more convenient way to communicate. In situations where typing is difficult, such as driving or when your hands are busy, this feature quickly converts voice into text, allowing users to transmit information without manual input. This significantly improves communication efficiency and meets users' needs for instant communication in a variety of complex scenarios.
[0003] However, when actually using the speech-to-text function, once entering the key reference analysis stage, if the user uses third-person pronouns, the system often finds it difficult to accurately distinguish "he", "she" and "it" based on the contextual semantic environment, which in turn causes semantic confusion and affects the accurate communication of information. Summary of the Invention
[0004] In view of this, the embodiments of the present disclosure provide a method, device, medium and product for key reference analysis in speech-to-text conversion, which can obtain the referent of third-person pronouns through semantic analysis, and combine the interpersonal relationship knowledge graph to match the referent in multiple dimensions to solve the semantic confusion problem caused by the difficulty in distinguishing third-person pronouns, thereby improving the quality of speech-to-text conversion and user experience.
[0005] In a first aspect, the embodiments of the present disclosure provide a method for parsing key references in speech-to-text conversion, which adopts the following technical solutions: Constructing a user's interpersonal relationship knowledge graph, and splitting the interpersonal relationship knowledge graph into multiple relationship sub-graphs according to interpersonal relationship types; wherein the interpersonal relationship types and the relationship sub-graphs each include a child node, and the child node includes a first annotation field and a second annotation field; Converting the user's speech into a text sequence, obtaining context information when the text sequence contains a third-person pronoun and the third-person pronoun meets a preset condition, and obtaining, based on the context information, the referent of the third-person pronoun, the type of relationship between the referent and the user, and the confidence level of the relationship type; Based on the relationship type and the confidence level, all relationship subgraphs are sorted to form a relationship subgraph set; Traverse the relationship sub-graph set in sorted order to obtain candidate relationship sub-graphs; Matching the referents with the child nodes and the first annotation fields of the child nodes in the candidate relationship subgraph respectively; When the referent successfully matches the child node or the referent successfully matches the first annotation field of the child node, a text sequence containing a correct third-person pronoun is obtained based on the second annotation field of the child node.
[0006] Optionally, constructing the user's interpersonal relationship knowledge graph includes: Obtain the user's contact note information, and build an initial knowledge graph based on the contact note information; Obtaining historical conversation data of the user on a preset chat platform, and extracting interpersonal relationship information from the historical conversation data; Based on the interpersonal relationship information, the initial knowledge graph is optimized to obtain a recommended knowledge graph; The recommended knowledge graph is displayed to the user, and in response to the user's modification operation instruction, the recommended knowledge graph is modified to generate an interpersonal relationship knowledge graph.
[0007] Optionally, obtaining the user's contact note information and constructing an initial knowledge graph based on the contact note information includes: Extracting name note information from the contact note information, and identifying whether the name note information contains preset characters; If the preset characters are included, extracting the primary title and the secondary title from the name note information based on the preset characters; If the preset characters are not included, the name note information is determined as the main title; The user's basic information is used as the central node, and the main title is used as the child node of the central node; Construct the first annotation field of the child node; If there is a secondary title corresponding to the child node, the secondary title is stored in the first annotation field of the child node; If there is no secondary title corresponding to the child node, the first annotation field of the child node is empty; Detecting the gender information included in the contact's remarks information, and storing the third-person pronoun corresponding to the gender information in the second annotation field of the child node; Create directed edges to represent the relationship between users and contacts; The central node, the child nodes, the first annotation field, the first annotation field and the directed edges constitute an initial knowledge graph.
[0008] Optionally, the acquiring of context information, based on the context information, acquiring the referent of the third-person pronoun, the type of relationship between the referent and the user, and the confidence level of the relationship type, includes: According to a preset first search range, short-distance context information is obtained with the sentence containing the third-person pronoun as a reference point, and a first prompt word is constructed based on the short-distance context information; If the AI model outputs the referent based on the first prompt word, obtaining the relationship type and the confidence level based on the short-distance context information and the AI model; If the AI model does not output the referent based on the first prompt word, obtaining long-distance context information based on the sentence containing the third-person pronoun according to a preset second search range and constructing a second prompt word based on the long-distance context information; If the AI model outputs the referent based on the second prompt word, obtaining the relationship type and the confidence level based on the long-distance context information and the AI model; If the AI model does not output the referent based on the second prompt word, obtaining cross-conversation context information based on a preset third search range and taking the timestamp of the sentence containing the third-person pronoun as a reference point, and constructing a third prompt word based on the cross-conversation context information; Based on the third prompt word and the AI model, the referent object, the relationship type and the confidence level are obtained.
[0009] Optionally, the acquiring of short-distance context information based on a preset first search range and taking a sentence containing a third-person pronoun as a reference point includes: Taking the sentence containing the third-person pronoun as the reference point, N sentences are taken forward and backward as short-distance context information.
[0010] Optionally, the acquiring of long-distance context information based on a preset second search range and taking a sentence containing a third-person pronoun as a reference point includes: Taking the sentence containing the third-person pronoun as the reference point, the conversation content consistent with the topic of the sentence is extracted forward and backward respectively, and the extracted conversation content is used as the long-distance context information.
[0011] Optionally, obtaining cross-dialogue context information according to a preset third search range and taking the timestamp of a sentence containing a third-person pronoun as a reference point includes: Taking the timestamp of the sentence containing the third-person pronoun as the reference point, set the first time interval forward and the second time interval backward; Cross-dialogue context information is filtered out within a time range formed by the first time interval and the second time interval.
[0012] Optionally, the key reference parsing method in the speech-to-text conversion further includes: According to the preset fourth search range, downstream information is obtained with the sentence containing the third person pronoun as a reference point; If the downstream information includes a preset prompt word, determining that the third-person pronoun does not meet the preset condition, and obtaining a text sequence including the correct third-person pronoun based on the preset prompt word; If the downstream information does not include the preset prompt word, it is determined that the third-person pronoun meets the preset condition.
[0013] In a second aspect, the embodiments of the present disclosure further provide a key reference parsing system in speech-to-text conversion, which adopts the following technical solutions: A knowledge graph construction module is used to construct a user's interpersonal relationship knowledge graph, and split the interpersonal relationship knowledge graph into multiple relationship sub-graphs according to interpersonal relationship types; wherein the interpersonal relationship types and the relationship sub-graphs each include a child node, and the child node includes a first annotation field and a second annotation field; A speech-to-text module is configured to convert the user's speech into a text sequence, obtain context information when the text sequence contains a third-person pronoun, and based on the context information, obtain the referent of the third-person pronoun, the type of relationship between the referent and the user, and the confidence level of the relationship type; A graph set forming module, configured to sort all relationship sub-graphs based on the relationship type and the confidence level to form a relationship sub-graph set; A candidate graph acquisition module is used to traverse the relationship sub-graph set in sorted order to obtain candidate relationship sub-graphs; A referent matching module, configured to match the referent with the child nodes and the first annotation fields of the child nodes in the candidate relationship subgraph respectively; The correct sequence acquisition module is used to acquire a text sequence containing a correct third-person pronoun based on the second annotation field of the child node when the reference object successfully matches the child node or the reference object successfully matches the first annotation field of the child node.
[0014] In a third aspect, the embodiments of the present disclosure further provide a computer device that adopts the following technical solution: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform any of the above-mentioned key reference parsing methods in speech-to-text conversion.
[0015] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute any of the above-mentioned key reference parsing methods in speech-to-text conversion.
[0016] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, comprising a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.
[0017] The key reference parsing method for speech-to-text conversion provided by the disclosed embodiments builds a user's interpersonal relationship knowledge graph, enabling the system to integrate and organize the various interpersonal relationships in the user's life. This provides a foundation for subsequent accurate parsing of third-person pronoun references, allowing the system to find suitable referents within a relatively complete interpersonal relationship system. By splitting the interpersonal relationship knowledge graph into multiple relationship subgraphs and dividing them according to interpersonal relationship types, the system can conduct more targeted searches when handling reference issues. Different relationship subgraphs correspond to different interpersonal relationship types, improving search efficiency and accuracy. Child nodes contain a first annotated field and a second annotated field. These annotated fields can store more detailed information about interpersonal relationships, such as a contact's secondary title and corresponding third-person pronouns. This helps the system make decisions based on richer information when matching referents, thereby improving matching accuracy. Converting the user's speech into a text sequence enables the system to perform text processing on the speech content, which is a prerequisite for subsequent reference parsing. Only by converting the speech into text can the third-person pronouns and their referent relationships be further analyzed. When a text sequence contains a third-person pronoun, the system can obtain contextual information and, based on this, determine the referent of the third-person pronoun, the type of relationship between the referent and the user, and the confidence level of the relationship type. This allows the system to extract key information from the text and provides a basis for subsequent sorting and matching. All relationship subgraphs are sorted based on relationship type and confidence level to form a set of relationship subgraphs. The sorting process allows the system to traverse the set of relationship subgraphs in sorted order, prioritizing relationship subgraphs with closer relationships to the referent and higher confidence levels. This narrows the search scope, reduces unnecessary search and matching processes, and improves search accuracy and efficiency. The referent is matched against the child nodes and the first annotated fields of the child nodes in the candidate relationship subgraphs. Through this multi-dimensional matching approach, the system can more accurately find the child nodes that match the referent, thereby determining the correct referential relationship. When the referent successfully matches the child node or the first annotation field of the child node, a text sequence containing the correct third-person pronoun is obtained based on the second annotation field of the child node. This step ultimately solves the semantic confusion problem caused by the difficulty in distinguishing third-person pronouns in the existing technology, ensuring the accurate transmission of information.
[0018] The above description is only an overview of the technical solution of the present disclosure. In order to more clearly understand the technical means of the present disclosure, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the following specifically cites preferred embodiments and describes them in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 A flowchart of a method for parsing key references in speech-to-text conversion provided by an embodiment of the present disclosure; Figure 2 A schematic diagram of a process for constructing a knowledge graph of interpersonal relationships according to an embodiment of the present disclosure; Figure 3 A flowchart of a method for obtaining a reference object, a relationship type, and a confidence level provided in an embodiment of the present disclosure; Figure 4 A flowchart of a method for resolving correct titles provided in an embodiment of the present disclosure; Figure 5 A flowchart of a method for constructing a vocabulary of rare Chinese characters in a title provided by an embodiment of the present disclosure; Figure 6 A flowchart of a method for obtaining a correct title provided in an embodiment of the present disclosure; Figure 7 A block diagram showing the principle of a key reference parsing system in speech-to-text conversion provided by an embodiment of the present disclosure; Figure 8 A schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0022] It should be clear that the following embodiments of the present disclosure are described through specific concrete examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other in the absence of conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.
[0023] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this device and / or practice this method.
[0024] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The illustrations only show components related to the present disclosure and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.
[0025] Additionally, in the following description, specific details are provided to provide a thorough understanding of the examples. However, one skilled in the art will appreciate that the aspects described can be practiced without these specific details.
[0026] Reference Figure 1 The present disclosure provides a method for parsing key references in speech-to-text conversion, comprising the following steps: S1: Construct a user's interpersonal relationship knowledge graph, and split the interpersonal relationship knowledge graph into multiple relationship sub-graphs according to interpersonal relationship types; wherein both the interpersonal relationship types and the relationship sub-graphs include child nodes, and the child nodes include a first annotation field and a second annotation field; S2: When the text sequence contains a third-person pronoun and the third-person pronoun meets a preset condition, context information is obtained. Based on the context information, the referent of the third-person pronoun, the relationship type between the referent and the user, and the confidence level of the relationship type are obtained. S3: Based on the relationship type and confidence, all relationship subgraphs are sorted to form a relationship subgraph set; S4: traverse the relationship subgraph set in sorted order to obtain candidate relationship subgraphs; S5: Match the referent with the child nodes in the candidate relationship subgraph and the first annotation field of the child nodes respectively; S6: When the referent matches the child node successfully or the referent matches the first annotation field of the child node successfully, a text sequence containing the correct third-person pronoun is obtained based on the second annotation field of the child node.
[0027] The key reference parsing method in speech-to-text provided by the present disclosure can integrate and sort out the various interpersonal relationships in the user's life by constructing the user's interpersonal relationship knowledge graph. This provides a basis for the subsequent accurate parsing of the reference of third-person pronouns, allowing the system to find suitable reference objects in a relatively complete interpersonal relationship system. The interpersonal relationship knowledge graph is split into multiple relationship sub-graphs and divided according to the type of interpersonal relationship, so that the system can search more specifically when dealing with reference problems. Different relationship sub-graphs correspond to different interpersonal relationship types, which improves the efficiency and accuracy of the search.
[0028] The child node contains a first annotation field and a second annotation field. These annotation fields can store more detailed information about interpersonal relationships, such as the contact's secondary title, corresponding third-person pronouns, etc., which helps the system make judgments based on richer information when matching the referent object, thereby improving the accuracy of the match.
[0029] Converting the user's speech into a text sequence enables the system to perform text processing on the speech content, which is a prerequisite for subsequent referential analysis. Only by converting the speech into text can the third-person pronouns and their referential relationships be further analyzed. When the text sequence contains third-person pronouns, the system can obtain contextual information and, based on this, obtain the referent of the third-person pronoun, the relationship type between the referent and the user, and the confidence of the relationship type. This enables the system to extract key information from the text and provide a basis for subsequent sorting and matching. All relationship subgraphs are sorted based on the relationship type and confidence to form a relationship subgraph set. The sorting process enables the system to traverse the relationship subgraph set in sorted order, thereby prioritizing relationship subgraphs with closer relationships to the referent and higher confidence. This can narrow the search scope, thereby reducing unnecessary search and matching processes and improving the accuracy and efficiency of the search.
[0030] The system matches the referent with the child nodes and the first annotated fields of the child nodes in the candidate relationship subgraph. This multi-dimensional matching approach allows the system to more accurately find the child nodes that match the referent, thereby determining the correct referential relationship. When the referent successfully matches the child node or the first annotated field of the child node, a text sequence containing the correct third-person pronoun is obtained based on the second annotated field of the child node. This step ultimately resolves the semantic confusion caused by the difficulty in distinguishing third-person pronouns in existing technologies, ensuring accurate communication of information.
[0031] In summary, by accurately parsing the referential relationship of third-person pronouns, the text sequence output by the system is more accurate and clear, which improves the quality of speech-to-text conversion. For users who rely on the speech-to-text function, this allows them to more accurately obtain and process information, thereby improving the user experience.
[0032] In S1, refer to Figure 2 The flowchart of the interpersonal relationship knowledge graph construction method is shown. "Constructing a user's interpersonal relationship knowledge graph" includes the following steps: S11: Obtain the user's contact note information and build an initial knowledge graph based on the contact note information; S12: Obtaining historical conversation data of the user on a preset chat platform, and extracting interpersonal relationship information from the historical conversation data; S13: Based on interpersonal relationship information, the initial knowledge graph is optimized to obtain a recommended knowledge graph; S14: Display the recommended knowledge graph to the user, and modify the recommended knowledge graph in response to the user's modification operation instructions to generate an interpersonal relationship knowledge graph.
[0033] In S11, the contents of the note fields of all contacts are extracted from the contact management system of the user device. The note information includes various information such as the contact's name, nickname, position, and relationship with the user.
[0034] Extract the name note information from the contact's note information and identify whether the name note information contains preset characters (such as brackets, quotation marks, dashes, etc.); if it contains preset characters, extract the primary and secondary titles from the name note information based on the preset characters; if it does not contain preset characters, determine the name note information as the primary title. For example, to identify the name note information containing preset characters, the note information in brackets and quotation marks is usually the secondary title, and the note information after the dash is also the secondary title. Therefore, if the name note information is Mr. Li (Lao Li), Mr. Li outside the brackets is the primary title, and Lao Li inside the brackets is the secondary title; if the name note information is "Lao Zhang" Manager Zhang, Lao Zhang inside the quotation marks is the secondary title, and Manager Zhang outside the quotation marks is the primary title; if the name note information is Boss Wang - Lao Wang, Boss Wang before the dash is the primary title, and Lao Wang after the dash is the secondary title. It can be seen that through this method, the extraction rules of the name note information can be determined according to the specific type of preset characters, and the main title and the secondary title can be extracted from the name note information based on the extraction rules.
[0035] The user's basic information is used as the central node, and the determined primary titles are used as subnodes of the central node. Each subnode represents a contact and is assigned a unique identifier in the system for subsequent management and query. The first annotation field of the subnode is constructed. If there is a secondary title corresponding to the subnode, the secondary title is stored in the first annotation field of the subnode; if there is no secondary title corresponding to the subnode, the first annotation field of the subnode is left empty.
[0036] Construct the second annotation field of the child node to detect whether the contact's remarks contain gender information; if so, store the third-person pronoun corresponding to the gender information in the second annotation field of the child node; if not, obtain the corresponding third-person pronoun of the contact by performing semantic analysis on the contact's remarks. If the acquisition fails, the second annotation field is first set to empty. If the acquisition is successful, store the corresponding third-person pronoun of the contact in the second annotation field of the child node. For example, if the contact's remarks contain "Ms. Zhang", when the system performs semantic analysis on the remarks, does the keyword "Ms." clearly point to a female gender? Therefore, the system can successfully obtain the third-person pronoun "she" corresponding to the contact and store "she" in the second annotation field of the child node.
[0037] Perform semantic analysis on contact notes. When the relationship between the user and the contact can be directly determined from the contact notes, create a directed edge to represent the relationship between the user and the contact. For example, if the user is female and one of the contact notes contains "Dad," a unidirectional edge is created from the central node to the "Dad" child node, labeled "Father-Daughter Relationship." If another contact note contains "Xiao Li" and "Colleague," a bidirectional edge is created between the central node and the "Xiao Li" child node, labeled "Colleague Relationship." If the relationship between the user and the contact cannot be determined from the contact notes, the edge between the central node and the contact is not created for now. The contact note is recorded and will be further refined in subsequent steps by combining historical conversation data and user modifications.
[0038] Through this method, it is possible to preliminarily obtain the user's interpersonal network based on the user's contact notes information, extract contacts associated with the user, and construct an initial knowledge graph based on this, that is, the central node, child nodes, first annotation field, second annotation field and directed edges constitute the initial knowledge graph. In this initial knowledge graph, it is possible to distinguish between the user's commonly used and uncommon names for contacts, use commonly used names as the main names to construct nodes, and use uncommon names as secondary names to enrich the node information, so as to facilitate the rapid matching of the user's voice-related information with the knowledge graph in actual use. Moreover, in actual use, the user's name for the same contact may be variable. In the existing technology, it is easy to ignore the user's nickname and other secondary names for the contact, resulting in a failure to match the knowledge graph, not fully utilizing the value of the knowledge graph, and thus affecting the accuracy of subsequent speech recognition. This method can effectively solve this problem by setting the first annotation field to store secondary names.
[0039] In S12, all historical conversation data of the user within a specified time period is obtained through an interface with a pre-defined chat platform (such as WeChat, QQ, etc.). Named entity recognition algorithms within natural language processing are used to identify person entities, such as names and titles, that appear in the historical conversation data. Conversation records between the user and each person entity are extracted from the historical conversation data. By analyzing the conversation records between the user and the person entity, relationship information between the user and the person entity is extracted. For example, if the sentence "Xiao Li and I completed this project together" appears in a conversation, it can be determined that "I" and "Xiao Li" are colleagues. The relationship information between the user and each person entity is combined to form the user's interpersonal relationship information.
[0040] In S13, the person entity in the interpersonal relationship information is compared with the child node and the first annotation field in the initial knowledge graph. If the person entity and the child node or the first annotation field of the person entity and the child node are the same, then based on the relationship information between the user and the person entity, the edge between the central node and the child node is optimized. For example, in the initial knowledge graph, some child nodes do not have edges. Through the relationship information between the user and the person entity, edges can be constructed for these nodes. Or, for edges that already exist in the initial knowledge graph, the relationship information between the user and the person entity can be used to find that the original edges have errors in pointing or annotation, and they need to be corrected. When analyzing the relationship information, for those situations where the relationship is consistent with that represented by the existing edges, the original edges are directly retained without any modification. If the first annotated fields of the person entity and the child node, and vice versa, are different, then it is determined whether the person entity and the child node refer to the same person. If so, the person entity is stored in the first annotated field of the child node. If not, the person entity is not included in the initial knowledge graph. A new child node is constructed in the initial knowledge graph based on the person entity. The relationship information between the user and each person entity is extracted from the interpersonal relationship information. Based on this relationship information, an edge is constructed between the central node and the new child node. Furthermore, if the second annotated field of the child node is empty, a semantic analysis of historical conversation data is performed to retrieve the corresponding third-person pronouns for the contact. If this retrieval fails, the second annotated field of the child node remains empty and can be supplemented by the user later. If this retrieval succeeds, the corresponding third-person pronouns for the contact are stored in the second annotated field of the child node. This method fully utilizes interpersonal relationship information to improve the initial knowledge graph, resulting in a user-friendly recommended knowledge graph.
[0041] In S14, the optimized recommended knowledge graph is visually displayed to the user, providing a user interface to allow the user to modify the knowledge graph, such as adding or deleting child nodes, modifying relationships, and supplementing the first and second annotation fields. Based on the user's modification instructions, the recommended knowledge graph is modified accordingly, ultimately generating an accurate interpersonal relationship knowledge graph.
[0042] Before splitting the interpersonal relationship knowledge graph, clarify the types of interpersonal relationships required, such as family relationships (such as parents, children, siblings, etc.), work relationships (such as colleagues, superiors, etc.), friendships, classmates, etc. These types can be further refined or merged based on specific business needs and data characteristics.
[0043] Traverse each child node and edge in the interpersonal relationship knowledge graph. For each edge, determine the type of interpersonal relationship it represents. Create an empty subgraph for each interpersonal relationship type, with the central node representing the user. When an edge belonging to a certain type is found, add the edge and its connected child nodes to the corresponding subgraph. For example, if an edge represents a parent-child relationship, add the edge and the corresponding parent-child node to the family relationship subgraph. Once the entire interpersonal relationship knowledge graph is traversed, multiple subgraphs are generated.
[0044] In S2, a microphone collects the user's voice signal and converts it into a digital signal. Speech features are then extracted and matched to acoustic patterns using pre-trained acoustic models (such as DNNs, RNNs, and their variants). A language model (statistical or based on deep learning) is then used to decode the phoneme sequence into a text sequence. Finally, post-processing is performed to ensure the output conforms to conventional expressions. During this processing, the text sequence is identified for third-person pronouns. If so, they are promptly modified to ensure the correct third-person pronouns are used.
[0045] Reference Figure 3 The flowchart of the method for obtaining the referent, relationship type, and confidence level is shown. The method of "obtaining context information and, based on the context information, obtaining the referent of the third-person pronoun, the relationship type between the referent and the user, and the confidence level of the relationship type" includes the following steps: S21: Acquire short-distance context information based on a preset first search range and taking a sentence containing a third-person pronoun as a reference point, and construct a first prompt word based on the short-distance context information; S22: Input the first prompt word into the AI model. If the AI model outputs a referent, execute S23; if the AI model does not output a referent, execute S24. S23: Obtain relationship type and confidence based on short-distance context information and AI model; S24: according to the preset second search range, taking the sentence containing the third-person pronoun as a reference point, obtaining long-distance context information, and constructing a second prompt word based on the long-distance context information; S25: Input the second prompt word into the AI model. If the AI model outputs a referent, execute S26; if the AI model does not output a referent, execute S27; S26: Obtain relationship type and confidence based on long-distance context information and AI models; S27: Acquire cross-dialogue context information according to a preset third search range, taking the timestamp of the sentence containing the third-person pronoun as a reference point, and construct a third prompt word based on the cross-dialogue context information; S28: Based on the third prompt word and the AI model, obtain the referent object, relationship type and confidence level.
[0046] In S21-S23, with the sentence containing the third-person pronoun as the reference point, N sentences are taken forward and backward as short-distance context information. The value of N can be preset according to the specific application scenario. For example, in an ordinary chat scenario, N can be set to 3-5 sentences. The reason for this setting is that within a short distance, the correlation between sentences is strong, and usually sufficient information can be provided to determine the referent of the third-person pronoun. Through this method, the sentence containing the third-person pronoun is located in the text sequence, and then according to the determined search range, the previous and next sentences can be extracted, the obtained short-distance context information is sorted, irrelevant symbols and modal particles are removed, and it is combined into a coherent text as the first prompt word.
[0047] The first prompt word is input into a pre-trained AI model. The AI model analyzes and processes the prompt word, attempting to output the referent of the third-person pronoun. If the AI model successfully outputs the referent, the process proceeds to S23. If not, it indicates insufficient short-range contextual information and requires expanding the search range, proceeding to S24.
[0048] If the AI model successfully outputs the referent, it will also analyze the relationship between the referent and the current conversation subject based on short-distance context information, such as friends, colleagues, family members, etc. The AI model can determine the relationship type by identifying keywords and character names in the context, and the AI model will give a confidence value to indicate the degree of certainty about the output relationship type.
[0049] In S24-S26, using sentences containing third-person pronouns as a reference point, the system traces back and forth to conversations consistent with the topic of the sentence. The extracted conversation content is used as long-range context information. This provides more background information and helps to more accurately determine the referent. Regarding conversation boundary characteristics, a conversation often revolves around one or more coherent topics. When there is a clear semantic break or a shift to a completely new, unrelated topic, it can be considered the end of the user's current conversation and the beginning of a new one. For example, in a business negotiation, if the price and specifications of a product are discussed first and then the topic suddenly shifts to weekend leisure activities, the different conversations can be divided based on the topic shift. Using sentences containing third-person pronouns as a reference point, the system traces back to the starting point that is semantically closely connected to the sentence and extends backward to the point where the semantic coherence ends. The content of this interval is used as long-range context information.
[0050] The acquired long-distance context information is organized, and irrelevant symbols and modal particles are removed to form a coherent text as the second prompt word. The second prompt word is input into the AI model, and the next step is determined based on the AI model's output. If the AI model outputs a referent, the process proceeds to S26; if not, the search scope needs to be further expanded, proceeding to S27. Similar to S23, if the AI model successfully outputs a referent, it also analyzes the relationship type between the referent and the current conversation subject based on the long-distance context information and provides a corresponding confidence value.
[0051] In S27 and S28, using the timestamp of the sentence containing the third-person pronoun as the reference point, a first time interval is set forward and a second time interval is set backward, for example, 24 hours before and 2 minutes after. Within the time range formed by the first and second time intervals, cross-conversation context information is filtered. This information includes the user's current conversation and other conversations related to the current conversation. For example, on a chat platform, suppose a conversation between the user and friend A at 3:30 PM on July 17, 2025, in which the user mentioned "he's planning a party." Using 3:30 PM as the reference point, cross-conversation information is filtered 24 hours forward and 2 minutes after. In a conversation between the user and friend B at 8:00 PM on July 16, B mentioned "C's planning a party this weekend." At 3:31 PM on July 17, friend D sent a message saying "I heard C's planning a party." These conversations are related to the current conversation mentioning "he" and together constitute cross-conversation context information, allowing us to infer that "he" likely refers to C.
[0052] Perform topic analysis on cross-conversation context information to filter out conversations that differ from the user's current conversation. Specifically, within a specified timeframe, filter all conversations other than the current one to remove those that are clearly unrelated to the current topic. For example, if the current conversation is about a work project, conversations related to entertainment and leisure can be excluded. This method extracts sentences related to the current topic, organizes the current conversation in the cross-conversation context information and the retained conversations that align with the current topic, and combines them into a coherent text, which serves as the third prompt word.
[0053] The third prompt word is input into the AI model. At this time, the third prompt word already contains a large amount of user conversation information and the content is comprehensive. The AI model can output the referent of the third-person pronoun, and accordingly, it can also output the relationship type between the referent and the current conversation subject and the corresponding confidence value.
[0054] The AI model consists of an embedding layer, a feature extraction layer, a referent prediction layer, and a relationship type determination layer. The embedding layer takes as input a prompt word (based on the corresponding prompt word input at different stages). The prompt word is in text form and constructed from contextual information obtained from different search scopes. The embedding layer converts the discrete prompt word into a continuous vector space representation to facilitate subsequent neural network processing. Using pre-trained word embedding models (such as Word2Vec and GloVe) or embedding layers fine-tuned for specific tasks, each word in the text is mapped into a fixed-length vector. The embedding layer converts each word in the prompt word into a low-dimensional vector representation, resulting in a sequence of word embedding vectors.
[0055] The sequence of word embedding vectors is input to the feature extraction layer, which processes the word embedding vectors using a deep learning model (such as a long short-term memory (LSTM) network, a gated recurrent unit (GRU) network, or a convolutional neural network (CNN) network) to extract contextual features of the text. These features reflect the relationships between words in the text and the overall semantics of the sentence. The feature vectors output by the feature extraction layer are fed to the referent prediction layer. Based on the extracted feature vectors, the referent prediction layer uses a fully connected layer or other classifier (such as a multi-layer perceptron (MLP)) to predict the referent. Furthermore, additional output nodes or regression layers are used to calculate the confidence of the referent prediction, which reflects the model's degree of certainty in the referent. If the confidence of any referent prediction is less than a preset threshold, the referent is considered undetermined and the referent prediction layer outputs either an empty string or a specific symbol to indicate that a prediction is undetermined. If the confidence of any referent prediction is greater than or equal to the preset threshold, the referent with the highest confidence is output.
[0056] The relationship type determination layer is activated only when the referent prediction is successful. It passes the feature vector output by the feature extraction layer and the referent identifier output by the referent prediction layer to the relationship type determination layer. After determining the referent, the relationship type determination layer further analyzes the relationship between the referent and the current conversation subject. A fully connected layer or other classifier processes the feature vector and referent information to determine the relationship type. Simultaneously, additional output nodes or regression layers are used to calculate the confidence level of the relationship type prediction. This method allows the relationship type determination layer to determine the relationship type between the referent and the current conversation subject, as well as the confidence level of the relationship type, if the referent prediction is successful. The confidence level ranges from [0 to 1].
[0057] Furthermore, the premise for obtaining the reference object, relationship type and confidence level mentioned above is that the third-person pronouns contained in the text sequence meet the preset conditions. The method for judging whether the preset conditions are met is: according to the preset fourth search range, the downstream information is obtained with the sentence containing the third-person pronoun as the reference point; if the downstream information contains the preset prompt word, it is determined that the third-person pronoun does not meet the preset conditions; if the downstream information does not contain the preset prompt word, it is determined that the third-person pronoun meets the preset conditions.
[0058] Downstream information refers to the P sentences obtained backward from the sentence containing the third-person pronoun as the reference point, where P≤N. If the downstream information contains preset prompt words, based on the preset prompt words, a text sequence containing the correct third-person pronoun is directly obtained. Preset prompt words include but are not limited to "she with a female radical," "he with a human radical," and "it with a treasure roof." For example, a user may say, "What's the child having for dinner tonight? Have you ordered takeout for her?" and then follow up with "she with a female radical." At this time, through voice recognition, the "he" with a male radical generated in the text sequence is automatically converted to "she" with a female radical.
[0059] In S3, the relationship subgraphs that match the relationship types are obtained and sorted in descending order of confidence to form an ordered set of relationship subgraphs. For example, suppose there are three known output relationship types: work, family, and friends, and there are also three corresponding relationship subgraphs. The confidence of the work relationship type is [X], the confidence of the family relationship type is [Y], and the confidence of the friend relationship type is [Z]. Sorting from highest to lowest confidence, if [X]>[Y]>[Z], the work relationship subgraph is ranked first, the family relationship subgraph is ranked second, and the friend relationship subgraph is ranked last, thus forming a set of relationship subgraphs.
[0060] In S4-S6, the relational subgraph set is traversed in the sorting order set in S3, and the relational subgraph currently being traversed is used as a candidate relational subgraph. The referent is first matched with the child node in the candidate relational subgraph; if the match with the child node is successful, the third-person pronoun stored in the second annotation field of the child node is used as the correct third-person pronoun, and the text sequence is corrected by the correct third-person pronoun to obtain a text sequence containing the correct third-person pronoun; if the match with the child node is unsuccessful, the referent is first matched with the first annotation field of the child node in the candidate relational subgraph; if the match with the first annotation field is successful, the third-person pronoun stored in the second annotation field of the child node is used as the correct third-person pronoun; if the match with the first annotation field is unsuccessful, return to step S4, continue to traverse the relational subgraph set, and obtain a new candidate relational subgraph until the match is successful or the relational subgraph set is traversed.
[0061] When the entire relationship subgraph is traversed and no match is found with the referent, the third-person pronoun used in the text sequence is determined to be correct and the text sequence is directly output. At the same time, the interpersonal relationship knowledge graph is updated based on the obtained referent, the relationship type between the referent and the user, and the third-person pronoun used in the text sequence.
[0062] Furthermore, in the speech recognition process, the problem of parsing rare characters in titles will be involved. Especially when processing names and other content containing rare characters, it is difficult for the system to accurately identify rare characters in titles, resulting in incorrect transcription results. Based on this, this application can solve the problem of inaccurate use of rare characters in the title parsing process by utilizing the constructed interpersonal relationship knowledge graph. Figure 4 The flowchart of the correct title resolution method is shown. The correct title resolution includes the following steps: S7: Build a user's rare word database; S8: When the text sequence includes a title, and the title is not a commonly used unique title, obtaining multiple title sequences; S9: When multiple appellation sequences contain uncommon characters, the correct appellation is determined from the multiple appellation sequences based on the interpersonal relationship knowledge graph and the appellation uncommon character lexicon; S10: Based on the correct title, obtain a text sequence containing the correct title sequence.
[0063] In S7, refer to Figure 5 The flowchart of the method for building a vocabulary of rare characters in titles is shown in the figure. "Building a vocabulary of rare characters in titles for users" includes the following steps: S71: extracting rare characters of titles from the standard character library, and constructing a basic word library based on the rare characters of titles; S72: Extracting the user's frequently used titles from the contact's notes and historical conversation data, optimizing the basic vocabulary based on the uncommon characters in the user's frequently used titles, and obtaining a vocabulary of uncommon characters in titles.
[0064] In S71, select an authoritative standard character library, such as the "Chinese Coded Character Set for Information Technology" released in 2022, which contains 87,900 characters. Filter out uncommon characters used for titles from the standard character library and organize the extracted uncommon characters into a vocabulary library as the basic vocabulary library.
[0065] In S72, analyze the contact note information, extract the name note information from it, and use natural language processing technology to extract appellation words from the historical conversation data. These name note information and appellation words are the commonly used appellations of the user. Extract rare Chinese characters from the commonly used appellations of the user. For example, the "Xun" in "Teacher Xun" is a rare Chinese character. Use these rare Chinese characters to optimize the basic vocabulary, remove the rare Chinese characters that are no longer used or have errors, and supplement the rare Chinese characters that the user actually uses but have not been stored. Finally, construct a rare Chinese character vocabulary for appellations. In addition, based on the historical conversation data, count the usage times of each rare Chinese character in the rare Chinese character vocabulary for appellations.
[0066] Adopt an incremental learning method to update the rare Chinese character vocabulary for appellations. Among them, with a six-month cycle, according to the changes in contact note information and the newly added conversation data, it is possible to perceive the changes in the user's usage habits, and then update the rare Chinese character vocabulary for appellations.
[0067] In S8, during the process of speech-to-text processing, when the text sequence contains an appellation, match the appellation with the pre-constructed commonly used appellation library to determine whether the appellation is a commonly used unique appellation. If it can be accurately matched and corresponds to a unique person in the actual context, it is a commonly used unique appellation, and the text sequence is directly output; if it is not in the commonly used appellation library or may correspond to multiple people, it is not a commonly used unique appellation, and multiple appellation sequences will be obtained for further clarification. For example, "Send a message to General Manager Jing", "Send a message to General Manager Jing", "Send a message to General Manager Jing", etc.
[0068] In S9, extract the different words between the appellations contained in multiple appellation sequences, and match the different words with the rare Chinese character vocabulary for appellations; if there are rare Chinese characters in the rare Chinese character vocabulary for appellations that are the same as the different words, it is determined that the multiple appellation sequences contain rare Chinese characters; if there are no rare Chinese characters in the rare Chinese character vocabulary for appellations that are the same as the different words, it is determined that the multiple appellation sequences do not contain rare Chinese characters. For example, for the three appellation sequences of "Send a message to General Manager Jing", "Send a message to General Manager Jing", and "Send a message to General Manager Jing", they respectively contain three appellations of "General Manager Jing", "General Manager Jing", and "General Manager Jing". Therefore, the different words are respectively "Jing", "Jing", and "Jing". As long as any one of the different words is included in the rare Chinese character vocabulary for appellations, it means that the multiple appellation sequences contain rare Chinese characters.
[0069] Refer to Figure 6 Referring to the flowchart of the correct appellation acquisition method shown, "Determine the correct appellation from multiple appellation sequences based on the interpersonal relationship knowledge graph and the rare Chinese character vocabulary for appellations" includes the following steps: S91: Based on the different words between the appellations contained in multiple appellation sequences, obtain the target syllables, and based on the target syllables, screen out the candidate rare Chinese characters with the same syllables from the rare Chinese character vocabulary for appellations and form a candidate rare Chinese character set; S92: Traverse the candidate rare Chinese character set in descending order of the number of uses, determine the current candidate rare Chinese character, and based on the current candidate rare Chinese character, obtain the current candidate title; S93: Obtain the voice termination time of the user, and based on the voice termination time and the preset time window size, extract from the user's current conversation the conversation fragment containing the title entity with the same pronunciation as the current candidate title; S94: Based on the current candidate title and the conversation fragment, determine the role attribution pointed to by the current candidate title, and based on the role attribution, determine the target relationship sub-graph from multiple relationship sub-graphs; S95: Detect whether the target relationship sub-graph contains the current candidate title; if it contains, execute S96; if it does not contain, execute S97; S96: Determine that the current candidate title is the correct title; S97: Continue to traverse the candidate rare Chinese character set, determine a new current candidate rare Chinese character until the correct title is obtained or the candidate rare Chinese character set traversal is completed.
[0070] In S91, use a pinyin input method or a pinyin annotation tool to obtain the syllables of the different words. For example, the syllables of "荆", "经", and "京" are "jing". The title rare Chinese character library adopts a storage method classified based on syllables and sorted by usage frequency. Classify all the rare Chinese characters in the title rare Chinese character library according to the syllables of the rare Chinese characters, group the rare Chinese characters with the same syllable into a set of rare Chinese characters, and use the same syllable as the index of the set of rare Chinese characters; within each set of rare Chinese characters, arrange them in descending order of the number of uses of the rare Chinese characters. Such storage is convenient for quickly locating rare Chinese characters according to syllables, and at the same time can preferentially display the more frequently used rare Chinese characters. Therefore, by comparing the target syllable with the index of the set of rare Chinese characters, select the set of rare Chinese characters with the same index as the target syllable as the candidate rare Chinese character set.
[0071] In S92, each rare Chinese character contained in the candidate rare Chinese character set is the candidate rare Chinese character at this time. These candidate rare Chinese characters have been pre-sorted according to the number of uses, and the candidate rare Chinese characters with more uses are arranged in the front. Starting from the first element of the sorted candidate rare Chinese character set, sequentially select the candidate rare Chinese characters as the current candidate rare Chinese character.
[0072] Replace the current candidate rare Chinese character at the position of the different word to generate the current candidate title. For example, if the current candidate rare Chinese character is "旌", and the original titles are "荆总", "经总", and "京总", then the current candidate title can be "旌总".
[0073] In S93, for example, the system records that the voice termination time of the user is 14:30:12 on March 18, 2025, and the preset time window size is 2 hours. Then the system determines that the search time range is from 12:30:12 to 14:30:12 on March 18, 2025. The system extracts the dialogue fragments containing the appellation entities (such as "Jing Gong", "Jing Zong", etc.) with the same pronunciation as the current candidate appellation (such as "Jing Zong") from the user's conversations during the period from 12:30:12 to 14:30:12 on March 18, 2025. These dialogue fragments obtained by using the cross-period conversation memory method include statements describing interpersonal relationships, related events, etc., which are convenient for subsequent use by the large language model.
[0074] In S94, obtain the manifestation form of the current candidate appellation, and match the manifestation form with each appellation structure form in the preset set of appellation structure forms, where these appellation structure forms are summarized according to common role appellation rules; if the manifestation form is the same as the appellation structure form, then use the role positioning corresponding to the appellation structure form as the role attribution pointed to by the current candidate appellation. For example, when the appellation structure form is "rare character + Gong", the corresponding role positioning is "work role - technical personnel". When the current candidate appellation is "Gang Gong", since its manifestation form is the same as the appellation structure form "surname + Gong", it can be determined that the role attribution pointed to by "Gang Gong" is technical personnel in the work role. The appellation structure form "rare character + Zong" corresponds to the role positioning of "work role - leader". When the current candidate appellation is "Jing Zong", its role attribution pointed to can be determined as the leader in the work role. The appellation structure form "rare character + Jie" corresponds to the role positioning of "social friend role - female friend or acquaintance". When the current candidate appellation is "Kuai Jie", its role attribution pointed to can be determined as the female friend in the friend role.
[0075] If the presentation format differs from the title structure, the candidate title and conversation fragment are fed into a large language model. The large language model extracts key elements from the conversation fragment, including time information, scene description, key events, and themes. After extracting these key elements, the large language model performs a preliminary screening of possible role assignments based on the time information and scene description. During work hours and in a work setting, the work role is prioritized. During leisure time and in a family or leisure setting, the family or friend roles are prioritized. For example, in an office setting during work hours, ambiguous titles are initially considered from the perspective of the work role. Key events and themes are combined to further narrow the scope of role assignments. If the conversation revolves around topics such as work tasks and business collaboration, the role assignment is more likely to be a work role; if it touches on family matters and caring for family members, the role assignment is more likely to be a family role; and if it discusses topics such as leisure activities and hobbies, the role assignment is likely to be a friend role. For example, in a conversation about project bidding, the relevant titles are likely to indicate a work role. The previous analysis results are revised and confirmed by referring to the direct role clues provided by related vocabulary. If the associated word explicitly mentions a role relationship, then this is an important basis for determining the role of the current candidate title. For example, if it mentions "communicating order details with customers", the associated word "customer" clearly defines the role involved.
[0076] According to the determined role attribution, a corresponding sub-graph is selected from the multiple relationship sub-graphs that have been split previously. For example, if the role attribution is determined to be a family elder, the family relationship sub-graph is selected as the target relationship sub-graph.
[0077] In S95-S97, all nodes in the target relationship subgraph are traversed to see whether there is a node that matches the current candidate title; if there is a matching node, it means that the target relationship subgraph contains the current candidate title, and the current candidate title is determined to be the correct title; if there is no matching node, all first annotation fields in the target relationship subgraph are traversed to see whether there is a first annotation field that matches the current candidate title; if there is a matching first annotation field, it means that the target relationship subgraph contains the current candidate title, and the current candidate title is determined to be the correct title; if there is no matching first annotation field, it means that the target relationship subgraph does not contain the current candidate title, and the current candidate title is determined to be not the correct title, and the candidate uncommon character set is returned to, and the next candidate uncommon character is selected as the new current candidate uncommon character, and steps S92-S97 are repeated until the correct title is determined or the entire candidate uncommon character set is traversed.
[0078] When all the candidate uncommon characters in the candidate uncommon character set have been traversed and the correct title still cannot be determined, the conversation fragments are combined and the large language model is further used to analyze the interpersonal relationships and possible identity characteristics implied by the title to determine the correct title, or the candidate uncommon characters with the highest number of uses in the candidate uncommon character set are directly selected to form the correct title.
[0079] In multiple appellation sequences, determining the results based solely on the rare characters given in these multiple appellation sequences is often not comprehensive enough. To solve this problem, this method selects a corresponding set of candidate rare characters through the target syllable. This set of candidate rare characters covers more rare characters that may appear, greatly enriching the selection range of rare characters. Based on the user's usage habits, the candidate rare characters are screened and verified in order from most to least number of times they are used. This can fully take into account the user's actual usage preferences and frequency, avoid missing those characters that are rare but appear more frequently in actual use, and use the verification of the interpersonal relationship knowledge graph to obtain more accurate and correct titles.
[0080] Furthermore, when multiple title sequences do not contain uncommon characters, for example, if there are two original title sequences, "Send a message to Mr. Yu" and "Send a message to Mr. Yu," title words are extracted from each title sequence, such as "Mr. Yu" and "Mr. Yu." The role type to which each title word refers is determined, such as the role type "work." Based on this role type, a specific relationship subgraph is selected from multiple relationship subgraphs, and the title words contained in the specified relationship subgraph are determined as the correct title.
[0081] In S10, the correct title is combined with any one of the multiple title sequences to form a text sequence containing the correct title sequence. For example, if the original title sequences are three, namely "Send a message to Mr. Jing", "Send a message to Mr. Jing", and "Send a message to Mr. Jing", and the correct title is "Mr. Jing", then the correct title sequence is "Send a message to Mr. Jing", thus obtaining a text sequence that correctly identifies the title.
[0082] In summary, this solution addresses the deficiencies of existing speech-to-text technologies in the recognition of rare characters in appellations and the recognition of third-person pronouns, achieving significant optimization and improvement. By combining the correct appellation sequence and the correct third-person pronouns, it is ultimately possible to obtain a text sequence that includes the correct third-person pronouns. In terms of the recognition of rare characters in appellations, this solution breaks through the limitations of the general character library and can comprehensively and accurately recognize rare characters and uncommon names. Different from the situation where existing tools rely on a general character library of about 30,000 Chinese characters and it is difficult to cover rare characters, this solution constructs a more abundant and comprehensive character library system to ensure that rare characters can also be accurately recognized. For example, for a rare character like "Jing" with extremely low utilization rate, this solution significantly improves the recognition rate and effectively solves the problem of low recognition rate in existing technologies. At the same time, this solution simplifies the memory correction process, eliminating the need to use the built-in input method and manually correct multiple times to remember, significantly improving the usage efficiency. Moreover, by associating specific contacts with character roles, it accurately outputs names containing rare characters, avoiding frequent manual corrections due to recognition errors and truly realizing the convenience of the speech-to-text function.
[0083] In terms of the recognition of third-person pronouns, this solution has strong context semantic analysis capabilities and can accurately distinguish between "he", "she", and "it" based on the context semantic environment of the conversation, effectively avoiding semantic confusion. In scenarios such as family communication, when the user uses "he" to refer to a family member, this solution can accurately determine the referent object, improving the recognition accuracy of personal pronouns. Taking a chat platform as an example, the recognition accuracy of the "she" scenario in existing technologies is relatively low. When the user describes non-human things such as pets, existing technologies usually use "he" instead of "it", while this solution can significantly improve this situation. Whether in different scenarios such as family communication, daily communication, or formal office work, it can provide accurate and clear text conversion results for users, ensuring the accurate transmission of information, meeting the communication requirements of formal communication and office collaboration scenarios, and truly realizing the efficient and accurate application of the speech-to-text function in various scenarios.
[0084] Refer to Figure 7 , this disclosure provides a key reference resolution system in speech-to-text, including: A knowledge graph construction module 101 for constructing the user's interpersonal relationship knowledge graph and splitting the interpersonal relationship knowledge graph into multiple relationship sub-graphs according to the types of interpersonal relationships; wherein, both the types of interpersonal relationships and the relationship sub-graphs include sub-nodes, and the sub-nodes include a first annotation field and a second annotation field; A speech-to-text conversion module 102 for converting the user's speech into a text sequence. When the text sequence contains a third-person pronoun, it obtains the context information, and based on the context information, obtains the referent object of the third-person pronoun, the relationship type between the referent object and the user, and the confidence level of the relationship type; A graph set forming module 103 is used to sort all relationship sub-graphs based on relationship type and confidence to form a relationship sub-graph set; The candidate graph acquisition module 104 is used to traverse the relationship sub-graph set in sorted order to obtain candidate relationship sub-graphs; The referent matching module 105 is used to match the referent with the child nodes and the first annotation fields of the child nodes in the candidate relationship sub-graph respectively; The correct sequence acquisition module 106 is configured to acquire a text sequence containing a correct third-person pronoun based on the second annotation field of the child node when the referent matches the child node successfully or the referent matches the first annotation field of the child node successfully.
[0085] The various variations and specific examples of the key reference parsing method in speech-to-text provided above are also applicable to the key reference parsing system in speech-to-text provided in the present disclosure. Through the above detailed description of the key reference parsing method in speech-to-text, those skilled in the art can clearly understand the implementation method of the key reference parsing system in speech-to-text. For the sake of brevity of the specification, it will not be described in detail here.
[0086] A computer device according to an embodiment of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc.
[0087] The processor can be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and can control other components in the computer device to perform desired functions. In one embodiment of the present disclosure, the processor is used to execute the computer-readable instructions stored in the memory, causing the computer device to perform all or part of the steps of the key reference resolution method in speech-to-text conversion described in various embodiments of the present disclosure.
[0088] Those skilled in the art should understand that in order to solve the technical problem of how to obtain a good user experience, this embodiment may also include well-known structures such as a communication bus and an interface, and these well-known structures should also be included in the scope of protection of this disclosure.
[0089] like Figure 8The present invention provides a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Figure 8 The computer device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0090] like Figure 8 As shown, a computer device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) or programs loaded from a storage device into a random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0091] Typically, the following devices can be connected to the I / O interface: input devices such as sensors or visual information acquisition devices; output devices such as display screens; storage devices such as tapes and hard disks; and communication devices. The communication device can allow the computer device to communicate with other devices (such as edge computing devices) wirelessly or by wire to exchange data. Figure 8 A computer device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0092] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by the processor, all or part of the steps of the key reference parsing method in the speech-to-text conversion embodiment of the present disclosure are executed.
[0093] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.
[0094] According to the computer-readable storage medium of the embodiment of the present disclosure, non-transitory computer-readable instructions are stored thereon. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the method of each embodiment of the present disclosure are executed.
[0095] The above-mentioned computer-readable storage media include, but are not limited to, optical storage media (e.g., CD-ROMs and DVDs), magneto-optical storage media (e.g., MOs), magnetic storage media (e.g., magnetic tapes or mobile hard disks), media with built-in rewritable non-volatile memory (e.g., memory cards), and media with built-in ROM (e.g., ROM cartridges).
[0096] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.
[0097] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0098] In the present disclosure, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. The block diagrams of the devices, devices, equipment, and systems involved in the present disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "including," "comprising," "having," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0099] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.
[0100] It should also be noted that in the system and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0101] Various changes, substitutions, and modifications may be made to the technology described herein without departing from the teachings defined by the appended claims. Moreover, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of things, means, methods, and actions described above. Currently existing or later developed processes, machines, manufactures, compositions of things, means, methods, or actions that perform substantially the same function or achieve substantially the same results as the corresponding aspects described herein may be utilized. Accordingly, the appended claims include within their scope such processes, machines, manufactures, compositions of things, means, methods, or actions.
[0102] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0103] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A method for analyzing key references in speech-to-text conversion, characterized in that: include: Constructing a user's interpersonal relationship knowledge graph, and splitting the interpersonal relationship knowledge graph into multiple relationship sub-graphs according to interpersonal relationship types; wherein the interpersonal relationship types and the relationship sub-graphs each include a child node, and the child node includes a first annotation field and a second annotation field; Converting the user's speech into a text sequence, obtaining context information when the text sequence contains a third-person pronoun and the third-person pronoun meets a preset condition, and obtaining, based on the context information, the referent of the third-person pronoun, the type of relationship between the referent and the user, and the confidence level of the relationship type; Based on the relationship type and the confidence level, all relationship subgraphs are sorted to form a relationship subgraph set; Traverse the relationship sub-graph set in sorted order to obtain candidate relationship sub-graphs; Matching the referents with the child nodes and the first annotation fields of the child nodes in the candidate relationship subgraph respectively; When the referent successfully matches the child node or the referent successfully matches the first annotation field of the child node, a text sequence containing a correct third-person pronoun is obtained based on the second annotation field of the child node.
2. The method for analyzing key references in speech-to-text conversion according to claim 1, wherein: The construction of the user's interpersonal relationship knowledge graph includes: Obtain the user's contact note information, and build an initial knowledge graph based on the contact note information; Obtaining historical conversation data of the user on a preset chat platform, and extracting interpersonal relationship information from the historical conversation data; Based on the interpersonal relationship information, the initial knowledge graph is optimized to obtain a recommended knowledge graph; The recommended knowledge graph is displayed to the user, and in response to the user's modification operation instruction, the recommended knowledge graph is modified to generate an interpersonal relationship knowledge graph.
3. The method for analyzing key references in speech-to-text conversion according to claim 2, wherein: The step of obtaining the user's contact note information and constructing an initial knowledge graph based on the contact note information includes: Extracting name note information from the contact note information, and identifying whether the name note information contains preset characters; If the preset characters are included, extracting the primary title and the secondary title from the name note information based on the preset characters; If the preset characters are not included, the name note information is determined as the main title; The user's basic information is used as the central node, and the main title is used as the child node of the central node; Construct the first annotation field of the child node; If there is a secondary title corresponding to the child node, the secondary title is stored in the first annotation field of the child node; If there is no secondary title corresponding to the child node, the first annotation field of the child node is empty; Detecting the gender information included in the contact's remarks information, and storing the third-person pronoun corresponding to the gender information in the second annotation field of the child node; Create directed edges to represent the relationship between users and contacts; The central node, the child nodes, the first annotation field, the first annotation field and the directed edges constitute an initial knowledge graph.
4. The method for analyzing key references in speech-to-text conversion according to claim 1, wherein: The acquiring of context information, based on the context information, acquiring the referent of the third-person pronoun, the type of relationship between the referent and the user, and the confidence level of the relationship type, includes: According to a preset first search range, short-distance context information is obtained with the sentence containing the third-person pronoun as a reference point, and a first prompt word is constructed based on the short-distance context information; If the AI model outputs the referent based on the first prompt word, obtaining the relationship type and the confidence level based on the short-distance context information and the AI model; If the AI model does not output the referent based on the first prompt word, obtaining long-distance context information based on the sentence containing the third-person pronoun according to a preset second search range and constructing a second prompt word based on the long-distance context information; If the AI model outputs the referent based on the second prompt word, obtaining the relationship type and the confidence level based on the long-distance context information and the AI model; If the AI model does not output the referent based on the second prompt word, obtaining cross-conversation context information based on a preset third search range and taking the timestamp of the sentence containing the third-person pronoun as a reference point, and constructing a third prompt word based on the cross-conversation context information; Based on the third prompt word and the AI model, the referent object, the relationship type and the confidence level are obtained.
5. The method for analyzing key references in speech-to-text conversion according to claim 4, characterized in that: The method of obtaining short-distance context information based on a preset first search range and taking a sentence containing a third-person pronoun as a reference point includes: Taking the sentence containing the third-person pronoun as the reference point, N sentences are taken forward and backward as short-distance context information.
6. The method for analyzing key references in speech-to-text conversion according to claim 4, wherein: The step of obtaining long-distance context information based on the preset second search range and taking the sentence containing the third-person pronoun as a reference point includes: Taking the sentence containing the third-person pronoun as the reference point, the conversation content consistent with the topic of the sentence is extracted forward and backward respectively, and the extracted conversation content is used as the long-distance context information.
7. The method for analyzing key references in speech-to-text conversion according to claim 4, wherein: The step of obtaining cross-dialogue context information based on the preset third search range and taking the timestamp of the sentence containing the third-person pronoun as a reference point includes: Taking the timestamp of the sentence containing the third-person pronoun as the reference point, set the first time interval forward and the second time interval backward; Cross-dialogue context information is filtered out within a time range formed by the first time interval and the second time interval.
8. The method for analyzing key references in speech-to-text conversion according to claim 4, wherein: Also includes: According to the preset fourth search range, downstream information is obtained with the sentence containing the third person pronoun as a reference point; If the downstream information includes a preset prompt word, determining that the third-person pronoun does not meet the preset condition, and obtaining a text sequence including the correct third-person pronoun based on the preset prompt word; If the downstream information does not include the preset prompt word, it is determined that the third-person pronoun meets the preset condition.
9. A computer device, characterized in that: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the key reference parsing method in speech-to-text conversion according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the key reference parsing method in speech-to-text conversion according to any one of claims 1 to 8.
11. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Intelligent call shorthand method and system, terminal and medium
CN121598908A