Conversation contradiction recognition method and system based on thinking chain of acoustic model and large language model

By combining acoustic models with large language model thinking chain technology, the speaker's identity is automatically identified and the dialogue content is analyzed, solving the problem of speaker role and contradiction identification in complex dialogue scenarios, and achieving more efficient information perception and decision support.

CN121884791APending Publication Date: 2026-04-17中国司法大数据研究院有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
中国司法大数据研究院有限公司
Filing Date
2025-12-03
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify speaker roles and points of conflict in complex dialogue scenarios, leading to low information perception efficiency and a high risk of decision-making errors.

Method used

By combining acoustic models with large language model thinking chain technology, the system automatically identifies the speaker's identity through speech recognition, speech segmentation, and role verification, and analyzes the core information of events in the dialogue context to quickly identify contradictions.

Benefits of technology

It significantly improves the speaker's information perception efficiency in complex dialogue scenarios, helps to quickly focus on key contradictions and issues, and reduces cognitive load and the risk of decision-making errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884791A_ABST
    Figure CN121884791A_ABST
Patent Text Reader

Abstract

The invention discloses a dialogue contradiction recognition method and system based on an acoustic model and a large language model thinking chain, and belongs to the field of artificial intelligence. The method comprises the following steps: processing mixed voice signals of multiple speakers to obtain clause text data of the speakers; according to the speaker clause timestamps, assembling the speaker clause text information and the corresponding speaker service role names to obtain dialogue context content; based on a private knowledge base, obtaining a knowledge entry text set of the dialogue context content; and on the basis of the dialogue context content and the knowledge item text set of the dialogue context content, calling a large language model to infer dialogue contradictions in the mixed voice signals of the multiple speakers. According to the invention, while the identity of the speaker can be automatically identified, the core information of the event in the dialogue content can be gradually mined, so that the contradiction point can be rapidly identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of speech recognition processing technology and natural language processing technology, specifically to a method and system for identifying dialogue contradictions based on the thought chain of acoustic models and large language models. Background Technology

[0002] In the current information technology environment, with the iteration of deep learning technology, especially the rapid development of speech recognition, voiceprint recognition, and large language model technology, the accuracy of speech recognition and the ability to understand natural language text have been greatly improved. This makes it possible to identify dialogue roles, understand dialogue content, and identify contradictions in dialogue through IT technology. To more accurately distinguish speaker roles, understand dialogue content more efficiently and intelligently, and more accurately identify potential contradictions in multi-party dialogues, this invention provides a method for constructing a core dialogue contradiction analysis engine. This method integrates acoustic model technology and large language model thinking chain technology to quickly and accurately identify voiceprint roles, analyze dialogue content, and identify contradictions in dialogue. This method can be flexibly applied to different scenarios such as judicial trials, dispute mediation, interrogation records, and corporate negotiations. It can significantly improve the information perception efficiency of speakers in complex dialogue scenarios, help them quickly focus on key contradictions, and reduce cognitive load and the risk of decision-making errors. Summary of the Invention

[0003] The purpose of this invention is to provide a method and system for dialogue contradiction identification based on acoustic models and a large language model thinking chain. It can automatically identify the speaker's identity based on acoustic models and role verification technologies, and analyze the semantics of each speaker in the dialogue context through a large model thinking chain to gradually extract the core information of the events in the dialogue content and quickly identify contradictions.

[0004] To achieve the above objectives, the technical solution of the present invention includes the following:

[0005] A dialogue contradiction recognition method based on the thought chain of acoustic models and large language models, the method comprising: The system processes mixed speech signals from multiple speakers to obtain the speaker's clause text data; wherein, the clause text data includes: speaker clause text information, speaker clause timestamp, and speaker's business role name; Based on the speaker clause timestamps, assemble the text information of each speaker clause and the corresponding speaker business role name to obtain the dialogue context content; Based on a private knowledge base, a set of knowledge entries texts related to the context of the dialogue is obtained; wherein, the knowledge data in the private knowledge base comes from laws and regulations, judicial practice, and industry standards. Based on the dialogue context and the set of knowledge entries for that dialogue context, a large language model is invoked to infer the dialogue contradictions in the mixed speech signals of the multiple speakers.

[0006] Furthermore, the mixed speech signals from multiple speakers are processed to obtain the sentence speech data of the target speaker, including: Speech streaming processing technology is used to process mixed speech signals from multiple speakers to obtain long speech streaming signals; Noise suppression techniques are used to denoise long speech streaming signals to obtain denoised speech signals. Speech endpoint detection technology is used to segment the denoised speech signal into phrases to obtain phrase speech signal groups; Role-based speech recognition technology is used to identify the speech signals of each clause in the clause speech signal group to obtain atomic speech data; wherein, the atomic speech data includes: atomic speech signal, atomic text, atomic timestamp and atomic role number; The atomic speech data are grouped according to the atomic role number to obtain speaker clause speech data; wherein, the speaker clause speech data includes: speaker clause speech signal, speaker clause text content and speaker clause timestamp; The speaker's clause speech signal is sent into the speech role verification sample pool. By calculating the similarity between the speaker's clause speech signal and the sample signal, the speaker's business role name is obtained. Generate clause text data that includes the speaker's clause text content, the speaker's clause timestamp, and the speaker's business role name.

[0007] Furthermore, the atomic speech data are grouped according to atomic role numbers to obtain speaker clause speech data, including: The atomic speech data are grouped according to the atomic role number to obtain several atomic speech data groups; For each atomic speech data group, the atomic speech signal and atomic text are concatenated in ascending order according to the atomic timestamp to obtain the speaker's clause speech signal and speaker's clause text content; The smallest atomic timestamp in the atomic speech data set is used as the speaker clause timestamp.

[0008] Furthermore, the private database includes: a text database and a vector database; The process of constructing the private knowledge base includes: Knowledge data is converted into text fragments with specific format characteristics to obtain knowledge entries; Identify and establish business relationships between knowledge items; Save the text content of knowledge entries and the business relationships between knowledge entries to a text database; The text content of knowledge entries is vectorized using an embedding model to obtain knowledge vector data, which is then saved to a vector database.

[0009] Furthermore, based on a private knowledge base, a set of knowledge entry texts related to the context of the dialogue is obtained, including: The context of the dialogue is segmented to obtain a set of tokens, and the text database is retrieved based on the tokens to obtain the first set of knowledge entry texts; The dialog context content is vectorized using an embedding model to obtain a set of context vector data. Based on this context vector data, a vector database is retrieved to obtain a second set of knowledge entry texts. Merge the first set of knowledge entries text with the second set of knowledge entries text to obtain the set of knowledge entries text for the context of the dialogue.

[0010] Furthermore, based on the dialogue context and the set of knowledge entries related to that context, a large language model is invoked to infer dialogue contradictions in the mixed speech signals of the multiple speakers, including: The first thought chain, aimed at identifying the events involved in the dialogue and extracting the core elements of the events, calls upon the large language model to analyze the context of the dialogue and obtain a list of events. The second thinking chain, aimed at identifying the correlation between events, calls the large language model to group the events in the event list, thus obtaining an event group list. The third thinking chain, aimed at identifying event contradictions, invokes a large language model to analyze the knowledge item text set of the event group list and the dialogue context content to obtain the contradiction identification results; By combining the contradiction identification results with the event list, the dialogue contradictions in the mixed speech signals of the multiple speakers can be obtained.

[0011] Furthermore, the prompt instruction template corresponding to the first thought chain includes dialogue text content slots and prompt corpus describing the task content, extraction rules, and output requirements; wherein, the extraction rules are used to define the core elements of the event extracted by the large language model, the core elements of the event include: event content, occurrence time, event process, event cause, event result, people, location, question and answer summary, and relationship attributes, the relationship attributes include: interpersonal relationships, interpersonal organizational relationships, and interpersonal asset relationships, the interpersonal relationships include: kinship relationships between people, job-related relationships between people, and legal relationships between people.

[0012] Furthermore, the prompt instruction template corresponding to the second thought chain includes an event list information slot and prompt corpus describing the task content and output requirements; wherein, the task content is to identify the same or similar events based on the descriptions of the event content field and the event cause field in the event list, group the events, reassemble the event object list by event group, and finally output it in JSON format; the event group information includes: event group ID, event group topic, earliest occurrence time and event object list.

[0013] Furthermore, the prompt instruction template corresponding to the third thinking chain includes an event group object list information slot, a reference knowledge information slot, and prompt corpus describing the task content, contradiction identification rules, and output requirements. The task content involves identifying fields with contradictory or conflicting meanings based on the details of each event in the event group object list, obtaining a list of conflicting events, and finally outputting it in JSON format. The contradictions or conflicts include: inconsistent descriptions of the same event by different people; inconsistent descriptions of the same event by the same person over and over; the time, place, and course of the event not conforming to common sense, folk customs, or universal values; the time, place, and course of the event not conforming to the descriptions of laws, legal provisions, and industry standards in the reference knowledge entries; and related contradictions between multiple events.

[0014] A dialogue contradiction recognition system based on an acoustic model and a large language model thought chain, the system comprising: The speech signal processing module is used to process mixed speech signals from multiple speakers and obtain the sentence text data of the speakers; wherein, the sentence text data includes: speaker sentence text information, speaker sentence timestamp, and speaker business role name; The context content generation module is used to assemble the text information of each speaker's clause and the corresponding speaker's business role name based on the speaker's clause timestamp to obtain the dialogue context content; The knowledge entry acquisition module is used to acquire a set of knowledge entry texts related to the context of the dialogue based on a private knowledge base; wherein, the knowledge data in the private knowledge base comes from laws and regulations, judicial practice, and industry standards. The dialogue contradiction generation module is used to infer dialogue contradictions in the mixed speech signals of multiple speakers by calling a large language model, based on the dialogue context content and the knowledge entry text set of the dialogue context content.

[0015] Compared with existing technologies, this invention integrates speech recognition technology, voiceprint recognition technology, and large language model technology to achieve more accurate speaker role recognition, more intelligent dialogue semantic understanding capabilities, and more reasonable and comprehensive dialogue contradiction recognition. It can significantly improve the speaker's information perception efficiency in complex dialogue scenarios, help them quickly focus on key contradictions, and reduce cognitive load and the risk of decision-making errors. Attached Figure Description

[0016] Figure 1 This is a flowchart of a dialogue contradiction recognition method based on the thought chain of acoustic models and large language models. Detailed Implementation

[0017] To further illustrate the technical solutions of the present invention, the present invention will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] This invention combines acoustic models and large language model thinking chain technology to identify the speaker's voiceprint features. It automatically identifies the speaker's identity through methods such as speech recognition, speech segmentation, speech fragmentation, and role verification. By analyzing the semantics of each speaker in the dialogue context through the large model thinking chain, it gradually extracts the core information of the events in the dialogue content and quickly identifies contradictions.

[0019] Figure 1 This is a flowchart of a dialogue contradiction recognition method based on the thought chain of acoustic models and large language models, which specifically includes the following steps.

[0020] Step 1: Process the mixed speech signals from multiple speakers to obtain the sentence text data of each speaker.

[0021] Step 1.1: In a multi-party voice dialogue scenario, the mixed voice signals of multiple speakers are continuously collected in real time, with each dialogue as a unit. The long voice streaming signal A is obtained by using voice streaming processing technology.

[0022] Step 1.2: Use noise suppression technology to denoise the speech signal A to obtain the denoised speech signal B.

[0023] Step 1.3: Using speech endpoint detection technology, the start and end times of effective speech in the long speech signal B are detected to achieve sentence segmentation and obtain the sentence speech signal group. .

[0024] Step 1.4: Employ role-based speech recognition technology to analyze the speech signals of each clause (e.g., The atomic speech data set was obtained through recognition. Each atomic speech data (such as (Contains atomic speech signals) Atomic speech text Atomic speech timestamp Atomic Speech Speaker Role Number Information such as the atomic speech speaker role number dRole and the clause speech signal (e.g., The units are incremented from 0 to 1.

[0025] Step 1.5: Employ a unique speech segmentation logic, based on the atomic speech speaker role number. The data in the atomic speech dataset is grouped to distinguish the speech data of different speakers. Within the same group, it is further categorized according to... Speech signals spliced ​​in ascending time order Get the speech signal of the current clause of the speaker. ,according to Concatenate the texts in ascending order of time Get the text content of the current clause of the speaker. Calculate the smallest atomic speech timestamp Get the current clause timestamp of the speaker. This allows us to obtain the speech data set of the speaker's clauses. Speech data for each speaker's clause (e.g.) )Include , , .

[0026] Step 1.6: Construct a voice role verification sample pool P, a pool named after the sample role number. Using sample speech signals as the key key-value pair arrays of values Sample Role Number The data is incremented from 0 to 1, based on the number of dialogues. Verification data will be added to the sample pool if either of the following conditions is met: 1. The sample pool is empty; 2. Voice role verification fails.

[0027] Step 1.7: Employ a unique voice role verification method to verify the speaker's clause voice data obtained in Step 1.5 (e.g., ...). The speech signal of the sentence obtained in step 1.5 is compared with the speech signals in the verification sample pool P one by one to calculate the voiceprint similarity, and finally determine the role of the sentence speech data. The speaker sentence speech data obtained in step 1.5 (such as...) The speech signal of the clause in ) As verification data, it is compared with the voice signals of each sample in the voice role verification sample pool P. Each voiceprint is fed into the speaker recognition model, where it is calculated using voiceprint recognition technology. If the similarity reaches or exceeds a threshold, the similarity is considered. If successful, return a confirmation message; otherwise, verification fails. Upon successful verification, assign a sample role number. As a clause role number Saved to speaker clause speech data (e.g.) In the process of verification failure, the verification data is treated as voice data for a new speaker role. Therefore, it is first processed in the largest sample role number in the current sample pool. Add 1 to the base to get a new sample role number. As a key, the verification voice signal will be used. As the value corresponding to this key, a new key-value pair is formed and added to the sample pool. Simultaneously, the new sample role number is assigned. As a clause role number Insert the speaker's clause speech data corresponding to the verification data (e.g.) This voice role verification method can significantly improve the success rate of speaker role recognition.

[0028] Step 1.8: Obtain the speaker's business role name corresponding to each role number in the sample pool P through user interaction. Such as judges, defendants, interrogators, and those being interrogated.

[0029] Step 1.9: Generate clause text data, which includes: speaker clause text information, speaker clause timestamp, and speaker business role name.

[0030] Step 2: Based on the speaker clause timestamps, assemble the text information of each speaker clause and the corresponding speaker business role name to obtain the dialogue context.

[0031] By clause timestamp Assemble the text information of each speaker's clause and the corresponding speaker's business role name in ascending order to obtain the dialogue context content F.

[0032] Step 3: Based on the private knowledge base, obtain the set of knowledge entries for the context of the dialogue.

[0033] Step 3.1: Construct a private knowledge base by collecting knowledge from laws and regulations, judicial practices, industry standards, etc. Based on the characteristics of the knowledge data format, use suitable technologies such as OCR to convert the knowledge data into text fragments K1 as knowledge entries. Identify and establish business relationships between knowledge entries, and save the text content of the knowledge entries and the relationships between them to a text database db1. Use an embedding model to vectorize the text content K1 of the knowledge entries to obtain knowledge vector data K2, and save the knowledge vector data K2 to a vector database db2.

[0034] Step 3.2: Segment the context content F to obtain a set of tokens. Search the text database db1 based on the tokens and obtain results with similarity scores higher than the threshold. Furthermore, the text content of the N knowledge entries with the highest similarity is extracted, thus obtaining a set of knowledge entry texts. The context content F is vectorized using an embedding model to obtain a set of context vector data. This data is then retrieved from the db2 vector database for similarity matching, and results are obtained when the similarity exceeds a threshold. The text content of the M knowledge entries with the highest similarity is obtained to form a knowledge entry text set. By merging the two sets, a text set of dialogue context-related knowledge entries can be obtained. .

[0035] Step 4: Based on the dialogue context content and the knowledge entry text set of the dialogue context content, call the large language model to infer the dialogue contradictions in the mixed speech signals of the multiple speakers.

[0036] Step 4.1: Analyze the obtained dialogue context content F to determine the events involved in the dialogue and extract the core elements of the events. Use the large model to group the events, obtaining an event list G. The prompt instruction template for the large language model includes dialogue text content slots and prompt corpora describing the task content, extraction rules, and output requirements. The core elements of the events include event content, time of occurrence, event process, cause of the event, event result, characters, location, question and answer summary, character relationships, and asset relationships. The text content slots are placeholders in the instruction template. The text in the dialogue context content F is concatenated using delimiters to form the following format and substituted into the text content slots, for example: [ {"Dialogue Role": "Judge", "Dialogue Text": "Plaintiff, please recall where you were and what you were doing at 8:00 AM on August 8, 2008."} {"Dialogue Character": "Plaintiff", "Dialogue Text": "I'm at a bar, watching the opening ceremony with the defendant..."}, {"Dialogue Character": "Judge", "Dialogue Text": "Defendant, please recall where you were and what you were doing at 8:00 AM on August 8, 2008."} {"Dialogue Character": "Defendant", "Dialogue Text": "The opening ceremony is at 8 PM. At 8 AM, I should be on my way to take my child to tutoring class. My child has class at 8:30 AM in Zhongguancun..."} ] The prompt instruction template corresponding to the first thought chain is as follows: # Task Description Please extract the speaker's business role field and event list from the following dialogue text according to the following extraction rules, and output them in JSON format.

[0037] # Extraction Rules - Event List: Individual event fields include event content, time of occurrence, event process, cause of event, event result, people involved, location, Q&A summary, and relationship fields (people to people, people to organizations, people to assets).

[0038] - The relational attributes are as follows: - Interpersonal relationships: - Extract kinship relationships between people (such as father and son, husband and wife, brothers and sisters, lovers, etc.).

[0039] - Extract the job relationships between people (such as superiors and subordinates, colleagues, partners, etc.).

[0040] - Extract the legal relationships between individuals (such as plaintiff and defendant, creditor and debtor, etc.).

[0041] - Relationship between people and organizations: - Extract the professional relationships (such as legal representatives, shareholders, employees, etc.) of the individuals and organizations involved in the incident.

[0042] - The relationship between people and assets: - Extract the ownership or guarantee relationship between the individual and the asset (such as the name of the asset owner, the type of asset owner (person / enterprise), the name of the asset, the value of the asset, the unit of value, the currency of value, etc.).

[0043] # Dialogue text content.

[0044] # The following is the output format { "Event List": [ { "Speaker's Business Role": (e.g., defendant) "Event Content": "A summary of the event, such as: Zhang San borrowing money from Li Si:", "Time of Occurrence": The time of the event, in a strict time format. ", "Cause of event": "Cause of event", "Event Details": "The specific details of the event must be completely consistent with the dialogue content." "Event Outcome": "Consequences of the event" "Event ID": "Event ID, a unique identifier for the event. Format requirements: Begins with 'event_', followed by 8 random uppercase and lowercase English characters and numbers." "Questions and Answers Summary": [ { "Question Summary": "A summary of the first question", "Answer Summary": A summary of the first answer. }, { "Question Summary": "A summary of the second question", "Answer Summary": "A summary of the second answer" }, { "Question Summary": "A summary of the Nth question", "Answer Summary": A summary of the Nth answer. } ], Original Q&A text: [ { Original Question Text: The original text of the first question. Original Answer: The original text of the first answer. }, { Original Question Text: The original text of the second question. Original answer: The original text of the second answer. }, { "Original Question Text": "The original text of the Nth question", "Original Answer": "The original text of the Nth answer" } ], "figure": [ { "Name": "If not specified or not clearly stated, this field will not be output". "ID Number": "If not available or not specified, this field will not be output". "Phone number": "If not available or not specified, this field will not be output". "Bank Card Number": "If not available or not specified, this field will not be output". "Work Unit": "If not specified or clearly stated, this field will not be output". "Job Title": "If none or not specified, this field will not be output." }, { "Name": "If not specified or not clearly stated, this field will not be output". "ID Number": "If not available or not specified, this field will not be output". "Phone number": "If not available or not specified, this field will not be output". "Bank Card Number": "If not available or not specified, this field will not be output". "Work Unit": "If not specified or clearly stated, this field will not be output". "Job Title": "If none or not specified, this field will not be output." } ], Location: "The location where the incident occurred" "Assets (items) involved": "Assets, items, etc. involved in the content", "relation": [ { "Relationship Type": "Person-to-Person" "Entity 1": { Entity type: "person" "name": "name", "Role": "Name of the marital relationship, such as: wife" }, Entity 2: { Entity type: "person" "name": "name", "Role": "Name of the spouse / husband relationship, such as: husband" }, Relationship Description: Marital Relationship }, { "Asset List": [ { Relationship Type: People and Assets "Entity 1": { Entity type: "person" "name": "name", "Role": "Affiliation, such as: Owner". }, Entity 2: { Entity Type: "Asset" "Asset Name": "Asset Name", "Asset Value": "Amount" "Value unit": "Yuan / Ten Thousand Yuan", "Currency of Value": "RMB, USD, HKD, etc." }, "Relationship Description": "Belonging Relationship" } ] } ] } ] } # Precautions - The large model strictly categorizes dialogue content according to context and semantics, and extracts event information strictly according to the speaker's expression.

[0045] - Strictly output each event as a separate JSON array, and the returned content cannot contain...

[0046] - A single conversation can contain multiple events, or even an unlimited number of events, and there is no limit to the number of events that can be extracted.

[0047] - A question summary is a summary of the speaker's question, while the original question text is the original text of the speaker's question. A response summary is a summary of the speaker's response, while the original response text is the original text of the speaker's response.

[0048] - The extracted results must be accurate and complete to avoid omitting key information.

[0049] Step 4.2: Based on the event list G, the second thought chain, aimed at identifying correlations between events, calls the large model to obtain the event group object list H. This involves constructing a prompt instruction template for the large language model, including event list information slots and prompt corpus describing the task content and output requirements. The event list information slots are placeholders in the instruction template, and the JSON format data of the event list obtained in step 12) is substituted into these slots. The prompt instruction template corresponding to the second thought chain is as follows: # Role You are a senior data annotation expert.

[0050] # Task Description Based on the descriptions in the event content and cause fields of the event list, identify identical or similar events, group the events, and reassemble the event object list by event group. Event group information includes: event group ID, event group topic, earliest occurrence time, and a list of event objects (one or more event objects), ultimately outputting in JSON format.

[0051] # The list of events to be entered.

[0052] # The following is the output format { "size": { "type": "number", "description": "Number of objects in the event group list" }, "List of event group objects": { "type": "array", "description": "List of event group objects", "items": { "type": "object", "description": "Event group object", "properties": { "Event Group ID": { "type": "string", "description": "Event group ID, a unique identifier for the event group. Format requirements: Begins with 'unionEvent_', followed by 6 random uppercase and lowercase English characters and numbers." }, "Event Group Topic": { "type": "string", "description": "Event group topic, summarizing the event name field of each event in the event object list." }, "Earliest occurrence time": { "type": "string", "description": "Earliest occurrence time: the earliest occurrence time among the events in this event group". }, "size": { "type": "number", "description": "Number of event objects" }, "List of event objects": { "type": "array", "description": "List of specific event objects", "items": { "type": "object", "description": "Single event object", "properties": { "Event ID": { "type": "string", "description": "Unique identifier ID of the event" }, "Time of occurrence": { "type": "string", "description": "The specific time the event occurred, in the format of year, month, day, hour, and minute." }, "Event Content": { "type": "string", "description": "The specific content or theme of the event" }, "Cause of incident": { "type": "string", "description": "A description of the cause of the event" }, "Event Summary": { "type": "string", "description": "A detailed description of the incident" }, "Event Outcome": { "type": "string", "description": "Description of the event outcome" }, Q&A Summary: { "type": "array", "description": "The content and format are completely consistent with the question-and-answer summary field in the input event object." } } } } } } } } # Require - Includes only field names and extracted field values: The generated JSON data only needs to contain the field names and their corresponding values, without needing type, description or other metadata information.

[0053] - Ensure structure matching: The field hierarchy and field names in the JSON must be consistent with the description structure.

[0054] - Handling missing fields: If a field is missing in the text, fill it with null or a reasonable default value and explain it in the output.

[0055] - No redundant nodes: Do not generate redundant keys or nodes, such as type.

[0056] - Output strictly according to JSON data format.

[0057] - The output list of event objects should include all fields of the original event object; you cannot output only the event ID.

[0058] Step 4.3: Based on the event group object list H and the dialogue context-related knowledge item set S, the third thought chain, aimed at identifying event contradictions, calls the large model to obtain the contradiction identification result J. Specifically, the prompt instruction template for constructing the large language model includes event group object list information slots, reference knowledge information slots, and prompt corpus describing the task content, contradiction identification rules, and output requirements. The event group object list information slots are placeholders in the instruction template; the obtained event group object list JSON format data is substituted into the event group object list information slots. The reference knowledge information slots are placeholders in the instruction template; the text information of each knowledge item in the obtained knowledge item text set S is converted into JSON format data and substituted into the reference knowledge information slots. The data format is as follows: [ "Text content of knowledge item 1", "Knowledge Item 2 Text Content", "Knowledge Item 3 Text Content" ] The conflict identification result J includes: conflict point title, conflict point analysis result, and conflict event information list. The conflict event information includes: unique event identifier, conflict field name, and conflict field content. The prompt instruction template corresponding to the third thought chain is as follows: # Role You are an experienced prosecutor, judge, police officer, member of the disciplinary inspection and review team, and also a senior data annotation expert.

[0059] # Task Description Based on the details of each event in the event group object list, identify fields with contradictory or conflicting meanings, obtain a list of conflicting events, and finally output them in JSON format.

[0060] Contradictions or conflicts include, but are not limited to, the following aspects: - Different people give inconsistent accounts of the same event.

[0061] - The same person's statements about the same event are inconsistent.

[0062] - The time, place, and course of the event do not conform to common sense, folk customs, or universal values.

[0063] - The time, place, and course of the event do not conform to the descriptions of laws, legal provisions, industry standards, etc. in the reference knowledge entries.

[0064] - There are contradictions among multiple events, such as conflicting or unreasonable timing of events, conflicting or unreasonable locations of events, inconsistent number / relationship of people, and unreasonable course of events.

[0065] The fields that need to be identified include: time of occurrence, location, name of the person, outcome of the event, cause of the event, course of the event, summary of the question, summary of the answer, relationship, and assets (items) involved.

[0066] Ignore fields include: event ID, event content, and original question and answer text.

[0067] # List of event group objects.

[0068] # Reference knowledge information.

[0069] # The following is the output format { "size": { "type": "number", "description": "Number of conflict points" }, "conflicts": { "type": "array", "description": "List of conflict points", "items": { "type": "object", "description": "Conflict point object", "properties": { "conflictTitle": { "type": "string", "description": "Conflict point title" }, "conflictDesc": { "type": "string", "description": "Conflict point analysis results" }, "eventSize": { "type": "number", "description": "Number of event objects" }, "events": { "type": "array", "description": "List of event objects", "items": { "type": "object", "description": "Single event object", "properties": { "eventId": { "type": "string", "description": "Unique identifier for the event" }, "fieldKey": { "type": "string", "description": "Name of the conflicting field" }, "fieldValue": { "type": "string", "description": "Content of the conflicting field" } } } } } } } } # Require - Includes only field names and extracted field values: The generated JSON data only needs to contain the field names and their corresponding values, without needing type, description or other metadata information.

[0070] - Ensure structure matching: The field hierarchy and field names in the JSON must be consistent with the description structure.

[0071] - Handling missing fields: If a field is missing in the text, fill it with null or a reasonable default value and explain it in the output.

[0072] - No redundant nodes: Do not generate redundant keys or nodes, such as type.

[0073] - Only outputs JSON data.

[0074] - Each event in the event object list needs to carry the original event unique identifier.

[0075] Step 4.4: Based on the unique event identifier in the contradiction identification result J, match it with the unique event identifiers in the obtained event list G to obtain all element information of the event, including the speaker's business role information, as supplementary information for the contradiction identification result J, thereby obtaining the final complete contradiction identification result information. The conflict identification results include: conflict point titles, conflict point analysis results, a list of conflict event information, information on each event element, and the speaker's role.

[0076] Although specific embodiments of the invention have been disclosed for illustrative purposes to aid in understanding and implementing the invention, those skilled in the art will understand that various substitutions, variations, and modifications are possible without departing from the spirit and scope of the invention and the appended claims. Therefore, the invention should not be limited to the content disclosed in the preferred embodiments, and the scope of protection claimed by the invention is defined by the claims.

Claims

1. A dialogue contradiction identification method based on an acoustic model and a large language model thinking chain, characterized by, The method includes: The system processes mixed speech signals from multiple speakers to obtain the speaker's clause text data; wherein, the clause text data includes: speaker clause text information, speaker clause timestamp, and speaker's business role name; Based on the speaker clause timestamps, assemble the text information of each speaker clause and the corresponding speaker business role name to obtain the dialogue context content; Based on a private knowledge base, a set of knowledge entries texts related to the context of the dialogue is obtained; wherein, the knowledge data in the private knowledge base comes from laws and regulations, judicial practice, and industry standards. Based on the dialogue context and the set of knowledge entries for that dialogue context, a large language model is invoked to infer the dialogue contradictions in the mixed speech signals of the multiple speakers.

2. The method of claim 1, wherein, Processing mixed speech signals from multiple speakers to obtain the sentence speech data of the target speaker, including: Speech streaming processing technology is used to process mixed speech signals from multiple speakers to obtain long speech streaming signals; Noise suppression techniques are used to denoise long speech streaming signals to obtain denoised speech signals. Speech endpoint detection technology is used to segment the denoised speech signal into phrases to obtain phrase speech signal groups; Role-based speech recognition technology is used to identify the speech signals of each clause in the clause speech signal group to obtain atomic speech data; wherein, the atomic speech data includes: atomic speech signal, atomic text, atomic timestamp and atomic role number; The atomic speech data are grouped according to the atomic role number to obtain speaker clause speech data; wherein, the speaker clause speech data includes: speaker clause speech signal, speaker clause text content and speaker clause timestamp; The speaker's clause speech signal is sent into the speech role verification sample pool. By calculating the similarity between the speaker's clause speech signal and the sample signal, the speaker's business role name is obtained. Generate clause text data that includes the speaker's clause text content, the speaker's clause timestamp, and the speaker's business role name.

3. The method according to claim 2, characterized in that, The atomic speech data is grouped according to the atomic role number to obtain the speaker clause speech data, including: The atomic speech data are grouped according to the atomic role number to obtain several atomic speech data groups; For each atomic speech data group, the atomic speech signal and atomic text are concatenated in ascending order according to the atomic timestamp to obtain the speaker's clause speech signal and speaker's clause text content; The smallest atomic timestamp in the atomic speech data set is used as the speaker clause timestamp.

4. The method according to claim 1, characterized in that, The private database includes: a text database and a vector database; The process of constructing the private knowledge base includes: Knowledge data is converted into text fragments with specific format characteristics to obtain knowledge entries; Identify and establish business relationships between knowledge items; Save the text content of knowledge entries and the business relationships between knowledge entries to a text database; The text content of knowledge entries is vectorized using an embedding model to obtain knowledge vector data, which is then saved to a vector database.

5. The method according to claim 4, characterized in that, Based on a private knowledge base, obtain the set of knowledge entries for the context of this dialogue, including: The context of the dialogue is segmented to obtain a set of tokens, and the text database is retrieved based on the tokens to obtain the first set of knowledge entry texts; The dialog context content is vectorized using an embedding model to obtain a set of context vector data. Based on this context vector data, a vector database is retrieved to obtain a set of second knowledge entry texts. Merge the first set of knowledge entries text with the second set of knowledge entries text to obtain the set of knowledge entries text for the context of the dialogue.

6. The method according to claim 1, characterized in that, Based on the dialogue context and the set of knowledge entries related to that context, a large language model is invoked to infer dialogue contradictions in the mixed speech signals of the multiple speakers, including: The first thought chain, aimed at identifying the events involved in the dialogue and extracting the core elements of the events, calls upon the large language model to analyze the context of the dialogue and obtain a list of events. The second thinking chain, aimed at identifying the correlation between events, calls the large language model to group the events in the event list, thus obtaining an event group list. The third thinking chain, aimed at identifying event contradictions, invokes a large language model to analyze the knowledge item text set of the event group list and the dialogue context content to obtain the contradiction identification results; By combining the contradiction identification results with the event list, the dialogue contradictions in the mixed speech signals of the multiple speakers can be obtained.

7. The method according to claim 6, characterized in that, The prompt instruction template corresponding to the first thought chain includes dialogue text content slots and prompt corpus describing the task content, extraction rules, and output requirements; wherein, the extraction rules are used to define the core elements of the event extracted by the large language model, and the core elements of the event include: event content, occurrence time, event process, event cause, event result, people, location, question and answer summary, and relationship attributes, and the relationship attributes include: interpersonal relationships, interpersonal relationships and inter-organizational relationships, and interpersonal relationships include: kinship relationships between people, job-related relationships between people, and legal relationships between people.

8. The method according to claim 6, characterized in that, The prompt instruction template corresponding to the second thinking chain includes an event list information slot and prompt corpus describing the task content and output requirements; wherein, the task content is to identify the same or similar events based on the descriptions of the event content field and the event cause field in the event list, group the events, reassemble the event object list by event group, and finally output it in JSON format; the event group information includes: event group ID, event group topic, earliest occurrence time and event object list.

9. The method according to claim 6, characterized in that, The prompt instruction template corresponding to the third thinking chain includes an event group object list information slot, a reference knowledge information slot, and prompt corpus describing the task content, contradiction identification rules, and output requirements. The task content involves identifying fields with contradictory or conflicting meanings based on the details of each event in the event group object list, obtaining a list of conflicting events, and finally outputting it in JSON format. The contradictions or conflicts include: inconsistent descriptions of the same event by different people; inconsistent descriptions of the same event by the same person over and over; the time, place, and course of the event not conforming to common sense, folk customs, or universal values; the time, place, and course of the event not conforming to the descriptions of laws, legal provisions, and industry standards in the reference knowledge entries; and related contradictions between multiple events.

10. A dialogue contradiction recognition system based on an acoustic model and a large language model thought chain, characterized in that, The system includes: The speech signal processing module is used to process mixed speech signals from multiple speakers and obtain the sentence text data of the speakers; wherein, the sentence text data includes: speaker sentence text information, speaker sentence timestamp, and speaker business role name; The context content generation module is used to assemble the text information of each speaker's clause and the corresponding speaker's business role name based on the speaker's clause timestamp to obtain the dialogue context content; The knowledge entry acquisition module is used to acquire a set of knowledge entry texts related to the context of the dialogue based on a private knowledge base; wherein, the knowledge data in the private knowledge base comes from laws and regulations, judicial practice, and industry standards. The dialogue contradiction generation module is used to infer dialogue contradictions in the mixed speech signals of multiple speakers by calling a large language model, based on the dialogue context content and the knowledge entry text set of the dialogue context content.

Citation Information

Patent Citations

  • Text semantic conflict detection method and system for constructing context association based on large model

    CN119474395A

  • Structured law document generation method and system based on large model and knowledge base enhancement

    CN120031008A

  • Intelligent dialogue system and method based on AI multi-mode large model

    CN120491834A

  • Acoustic event triggering-based agent real-time verbal skill recommendation method and system

    CN120913563A

  • Adaptive Virtual Intelligent Agent

    US20150142704A1