An AI-based intelligent question-answering scenario generation method and system

By collecting question-and-answer execution signals to generate answer recordings, identifying keywords, and asking follow-up questions about missing information, the problem of information omission caused by incomplete answers is solved, thus improving the efficiency and accuracy of question-and-answer interaction.

CN120611725BActive Publication Date: 2026-01-02杭州威灿科技有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510759627.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2026-01-02
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

In existing intelligent question-answering scenario generation technologies, incomplete answers lead to missing information, resulting in decreased interaction efficiency and failing to meet requirements.

Method used

By collecting question-and-answer execution signals, generating answer recordings and identifying content keywords, matching question-and-answer scenarios, generating baseline answer content, and asking follow-up questions for missing information, a complete transcript is generated.

Benefits of technology

Ensure no key information is omitted, improve the efficiency of question-and-answer interaction, generate complete transcripts with rigorous logic and sufficient evidence, and improve interaction efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611725B_ABST
    Figure CN120611725B_ABST
Patent Text Reader

Abstract

The application relates to an AI-based intelligent question and answer scene generation method and system, and relates to the field of natural voice processing, which comprises the following steps: collecting the answering audio of a person being asked in response to a question and answer execution signal; generating answering content and content keywords based on the answering audio; matching a question and answer scene based on the content keywords; generating reference answering content and a final questioning process in response to the question and answer scene; when the answering content is inconsistent with the reference answering content, obtaining missing answering content according to the answering content and the reference answering content; generating a follow-up question based on the missing answering content, and performing follow-up questioning; when the answering content is consistent with the reference answering content or after the follow-up questioning is completed, the questioning process continues to be prompted, and the current questioning process is collected; when the current questioning process is consistent with the final questioning process, the whole-scene answering audio is collected, the complete transcript is generated based on the whole-scene answering audio, and the complete transcript is uploaded to a transcript storage terminal. The application has the effect of improving the efficiency of scene interaction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of natural language processing, and in particular to an AI-based intelligent question and answer scenario generation method and system. BACKGROUND

[0002] Intelligent question and answer scenario generation refers to automatically constructing a question and answer interactive environment with specific scenario attributes, role relationships and dialogue logic using artificial intelligence technology.

[0003] Currently, intelligent question and answer scenario generation technology mainly relies on pre-set templates, fixed rule bases or statistical machine learning models to realize question and answer process automation. After the person being asked completes the answer, the content of the answer of the person being asked is recorded and subsequent questions are asked.

[0004] However, if the content of the answer of the person being asked is incomplete, key information may be missing, making it difficult to achieve the desired goal of question and answer interaction, thereby reducing the efficiency of scenario interaction, which needs to be improved. SUMMARY

[0005] In order to improve the efficiency of scenario interaction, the present application provides an AI-based intelligent question and answer scenario generation method and system.

[0006] In a first aspect, the present application provides an AI-based intelligent question and answer scenario generation method, which adopts the following technical solution:

[0007] An AI-based intelligent question and answer scenario generation method, comprising:

[0008] S1: collecting a question and answer execution signal of a pre-set question and answer terminal;

[0009] S2: collecting an answer recording of a person being asked in response to the question and answer execution signal;

[0010] S3: generating an answer content and a content keyword based on the answer recording;

[0011] S4: matching a question and answer scenario based on the content keyword;

[0012] S5: generating a reference answer content and a final question process in response to the question and answer scenario;

[0013] S50: when the answer content is inconsistent with the reference answer content, obtaining an answer missing content according to the answer content and the reference answer content;

[0014] S51: generating a follow-up question based on the answer missing content, and asking the follow-up question based on the follow-up question;

[0015] S52: When the answer content is consistent with the reference answer content or the follow-up question is completed, the question process continues to prompt and collect the current question process;

[0016] S53: When the current question process is consistent with the final question process, collect the whole scene answer recording, generate a complete record based on the whole scene answer recording, and upload the complete record to a preset record storage terminal.

[0017] By adopting the above technical solution, the question and answer execution signal is collected first, and the answer recording of the person being asked is obtained, the answer content and keywords are generated by voice recognition technology, and the question and answer scene is accurately positioned. After generating the reference answer content and the final question process based on the scene, the actual answer content is compared with the reference content, if there is inconsistency, the missing information is automatically extracted and the follow-up question is generated, to ensure that no key information is missed. When the answer content meets the standard or the follow-up question is completed, the system continues to promote the question process, until the current process is consistent with the final process, a complete record is automatically generated and uploaded and stored, so as to effectively solve the problem of missing information caused by incomplete answer content, and improve the efficiency of scene interaction.

[0018] Optionally, it also includes a complete record generation method:

[0019] S530: Generating a whole scene answer content based on the whole scene answer recording;

[0020] S531: Generating an answer content point in response to the whole scene answer content;

[0021] S532: Obtaining a content point filling position according to the answer content point and a preset question list;

[0022] S533: Filling the answer content point into the question list in the content point filling position to obtain a complete question list;

[0023] S534: Generating a visual evidence and a judgment result based on the complete question list, intelligently asking questions based on the judgment result, and collecting a judgment answer content;

[0024] S535: When the judgment answer content is consistent with the preset reference answer content, obtaining a complete record based on the complete question list, the visual evidence, the judgment result and the judgment answer content.

[0025] Optionally, it also includes a step after collecting the question and answer execution signal of the preset question and answer terminal:

[0026] S10: Collecting a question and answer terminal number in response to the question and answer execution signal;

[0027] S11: generating a terminal use scenario based on the question and answer terminal number;

[0028] S12: generating a recording collection method in response to the terminal use scenario;

[0029] S13: collecting an answer recording of the person being questioned based on the recording collection method.

[0030] Optionally, the recording collection method comprises:

[0031] S130: when the terminal use scenario is a preset indoor scenario, collecting indoor image information;

[0032] S1300: performing portrait posture recognition on the indoor image information to obtain a current speaking orientation of the person;

[0033] S1301: when the current speaking orientation of the person is inconsistent with a preset reference speaking orientation of the person, matching an auxiliary sound collection device number according to the current speaking orientation of the person;

[0034] S1302: based on the auxiliary sound collection device number, controlling a corresponding auxiliary sound collection device to be turned on, and collecting an auxiliary collection recording of the auxiliary sound collection device and a main collection recording of a preset main sound collection device;

[0035] S1303: obtaining an answer recording according to the auxiliary collection recording and the main collection recording.

[0036] Optionally, the method further comprises a verification method for the answer recording:

[0037] S13030: matching a demand video recording device number according to the current speaking orientation of the person;

[0038] S13031: based on the demand video recording device number, controlling a corresponding video recording device to be turned on, and collecting a speaking video of the person;

[0039] S13032: performing lip shape recognition on the person in the speaking video to obtain a lip shape action of the person;

[0040] S13033: obtaining an answer recording based on the lip shape action of the person, the auxiliary collection recording, and the main collection recording.

[0041] Optionally, the recording collection method further comprises:

[0042] S131: when the terminal use scenario is a preset outdoor scenario, collecting a recording signal of a preset outdoor sound collection device;

[0043] S1310: When the recording signal is consistent with the preset recording adjustment signal, adjusting the orientation of the person based on the recording signal;

[0044] S1311: Adjusting the orientation of the person to match the specific opening of the window;

[0045] S1312: Based on the specific opening of the window to obtain the vehicle recording device number;

[0046] S1313: Based on the vehicle recording device number to control the corresponding vehicle recording device to start, open the window corresponding to the specific opening of the window, and control the preset recording masking device to mask the vehicle recording device;

[0047] S1314: Collecting the vehicle collection recording of the vehicle recording device and the outdoor collection recording of the outdoor recording device;

[0048] S1315: Based on the vehicle collection recording and the outdoor collection recording to obtain the answer recording.

[0049] Optionally, the method also includes generating the answer content:

[0050] S30: Based on the answer recording to generate a recording text;

[0051] S31: Controlling the preset question and answer terminal to perform sentence detection on the recording text to obtain a sentence fluency detection result;

[0052] S32: Determining whether the sentence fluency detection result contains a preset sentence stuttering feature;

[0053] S320: When the sentence fluency detection result contains the sentence stuttering feature, performing stuttering processing with a preset stuttering processing method;

[0054] S321: When the sentence fluency detection result does not contain the sentence stuttering feature, based on the recording text to obtain the answer content.

[0055] Optionally, the stuttering processing method includes:

[0056] S3200: Based on the sentence fluency detection result to retrieve a stuttering sentence;

[0057] S3201: Based on the recording text and the stuttering sentence to obtain a stuttering timeline;

[0058] S3202: Collecting an answer recording video;

[0059] S3203: Based on the answer recording video and the stuttering timeline to determine a stuttering segment;

[0060] S3204: mouth shape recognition is performed on the character in the card fragment to obtain a character card mouth shape, and based on the character card mouth shape and the recording text, specific card content is obtained;

[0061] S3205: based on the card sentence and the specific card content, a smooth sentence is obtained.

[0062] Optionally, further comprising:

[0063] S60: collecting an answer recording video;

[0064] S61: mouth shape recognition is performed on the character in the answer recording video to obtain a video character mouth shape;

[0065] S62: based on the video character mouth shape, video text content is obtained;

[0066] S63: when and only when the video text content contains a preset prohibited word, a prohibited word abnormality prompt is reported.

[0067] In a second aspect, the present application provides an AI-based intelligent question and answer scenario generation system, which adopts the following technical solution:

[0068] An AI-based intelligent question and answer scenario generation system, comprising:

[0069] A collection module for collecting question and answer execution signals, answer recordings, current questioning processes, and complete answer recordings;

[0070] A memory for storing the program of any of the AI-based intelligent question and answer scenario generation methods;

[0071] A processor for loading and executing the program stored in the memory.

[0072] In summary, the present application includes the following at least one beneficial technical effect:

[0073] 1. First, the question and answer execution signals are collected and the answer recordings of the questioned person are obtained, the answer content and keywords are generated through voice recognition technology, and the question and answer scenario is accurately positioned. After generating the benchmark answer content and the final questioning process based on the scenario, the actual answer content is compared with the benchmark content, if there is inconsistency, the missing information is automatically extracted and follow-up questions are generated, ensuring that no key information is missed. When the answer content meets the requirements or the follow-up questions are completed, the system continues to advance the questioning process until the current process and the final process are consistent, and a complete record is automatically generated and uploaded and stored, thereby effectively solving the problem of missing information caused by incomplete answer content and improving the interaction efficiency;

[0074] 2. By first generating complete text content based on the entire answer recording, and further extracting key answer content points, through intelligent matching with the question list, accurately positioning the filling position of each content point, forming a structured complete question list. On this basis, the system automatically generates visual evidence and judgment results, and based on the judgment results, intelligent questioning is carried out to further verify the completeness and accuracy of the answers. When the judgment answer content is consistent with the reference answer content, the system will organically integrate the complete question list, visual evidence, judgment results and judgment answer content to form a complete record with rigorous logic and sufficient evidence. In this way, not only the comprehensiveness and accuracy of the record content are ensured, but also the authority and traceability of the record are improved through visual evidence and intelligent judgment;

[0075] 3. When the terminal use scene is an outdoor scene, first collect the recording signal of the outdoor recording equipment, if the signal is consistent with the recording adjustment signal, further analyze the personnel adjustment direction, accurately position the specific opened window. Through the vehicle window information matching corresponding vehicle recording equipment number, control the equipment to open and synchronously open the corresponding window, at the same time enable the recording shielding equipment to protect the vehicle recording equipment, effectively reduce the external interference while avoiding being detected. After synchronously collecting the vehicle collection recording and the outdoor collection recording by the double equipment, the system carries out intelligent fusion processing on the two audio signals to generate the final answer recording, which provides a solid and reliable data foundation for the subsequent question and answer analysis and record generation. BRIEF DESCRIPTION OF DRAWINGS

[0076] Figure 1 is a method flowchart of an AI-based intelligent question and answer scene generation method in an embodiment of the present application;

[0077] Figure 2 is a method flowchart of a complete record generation method in an embodiment of the present application;

[0078] Figure 3 is a step flowchart of the steps after collecting the pre-set question and answer terminal question and answer execution signal in an embodiment of the present application;

[0079] Figure 4 is a method flowchart of one of the recording collection methods in an embodiment of the present application;

[0080] Figure 5 is a method flowchart of an answer recording verification method in an embodiment of the present application;

[0081] Figure 6 is a method flowchart of another recording collection method in an embodiment of the present application;

[0082] Figure 7 is a method flowchart of an answer content generation method in an embodiment of the present application;

[0083] Figure 8 is a method flowchart of the method of the gop processing method in the embodiment of the application. DETAILED DESCRIPTION

[0084] The application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0085] Reference Figure 1 The embodiment of the present application discloses an AI-based intelligent question and answer scenario generation method, comprising the following steps:

[0086] S1: Collect the question and answer execution signal of the pre-set question and answer terminal.

[0087] The question and answer terminal refers to a terminal used to ask questions to the person being asked. In the present embodiment, the question and answer terminal is a mobile phone. The question and answer execution signal refers to a signal that needs to use the question and answer terminal to ask questions to the person being asked. The question and answer execution signal is obtained by calling the question and answer terminal. When it is necessary to ask questions to the person being asked, the question and answer terminal will issue a question and answer execution signal.

[0088] S2: Respond to the question and answer execution signal to collect the answer recording of the person being asked.

[0089] The answer recording refers to the recording of the person being asked when answering in the question and answer process. The answer recording is obtained by recording the question and answer terminal. When the question and answer execution signal is issued, the answer recording of the person being asked needs to be collected for subsequent steps.

[0090] S3: Based on the answer recording, generate the answer content and the content keyword.

[0091] The answer content refers to the text information converted from the answer recording of the person being asked. The content keyword refers to the key word, phrase or term extracted from the answer content. The answer recording can be converted into text content by using speech recognition technology, and then the answer content can be obtained. The content keyword in the answer content can be obtained by using natural language processing technology. Speech recognition technology and natural language processing technology are both well-known in the art, and will not be described here.

[0092] S4: Based on the content keyword, match the question and answer scenario.

[0093] The question and answer scenario refers to the specific business field involved in this question and answer. The question and answer scenario corresponding to the content keyword can be queried through the pre-set scenario comparison table, which records different question and answer scenarios corresponding to different content keywords. The scenario comparison table is formed by sequentially recording different question and answer scenarios corresponding to different content keywords by those skilled in the art, and will not be described here.

[0094] S5: Respond to the question and answer scenario to generate the reference answer content and the final question and answer process.

[0095] The reference answer content refers to the standardized reference answer content that the person being asked needs to make for a specific question and answer scene. The final questioning process refers to the last process of the person being asked when conducting the question and answer. The reference answer content and the final questioning process corresponding to the question and answer scene can be matched through a preset answer database, which stores different reference answer contents and final questioning processes corresponding to different question and answer scenes. The answer database is formed by recording the different reference answer contents and final questioning processes corresponding to different question and answer scenes in sequence by the person skilled in the art, which will not be repeated here.

[0096] S50: When the answer content is inconsistent with the reference answer content, the missing answer content is obtained according to the answer content and the reference answer content.

[0097] The missing answer content refers to the key information point that the answer content of the person being asked is missing compared with the reference answer content. By comparing the answer content and the reference answer content, the content in the reference answer content that is more than the answer content is known, and this part of the content is extracted to obtain the missing answer content.

[0098] When the answer content is inconsistent with the reference answer content, it means that the answer content of the person being asked is missing, and the missing answer content needs to be obtained first for the subsequent steps.

[0099] S51: Based on the missing answer content, a follow-up question is generated, and the follow-up question is asked based on the follow-up question.

[0100] The follow-up question refers to a supplementary question designed for the missing answer content. The follow-up question corresponding to the missing answer content can be queried through a preset question comparison table, which records the follow-up questions corresponding to different missing answer contents. The question comparison table is formed by recording the follow-up questions corresponding to different missing answer contents in sequence by the person skilled in the art, which will not be repeated here. At the same time, the question and answer terminal is controlled to ask the person being asked the follow-up question.

[0101] S52: When the answer content is consistent with the reference answer content or after the follow-up is completed, the questioning process continues to prompt and collects the current questioning process.

[0102] The current questioning process refers to the specific step currently conducted in the entire question and answer process. It can be queried through the question and answer terminal. The current questioning process is recorded in the question and answer terminal.

[0103] When the answer content is consistent with the reference answer content or after the follow-up is completed, it means that the question and answer of this process have been completed, and the questioning process continues to prompt is directly reported, and then the process continues, and the current questioning process is collected for the subsequent steps.

[0104] S53: Collect the whole answering recording based on the current questioning procedure being consistent with the final questioning procedure, generate a complete transcript based on the whole answering recording, and upload the complete transcript to a preset transcript storage terminal.

[0105] The whole answering recording refers to the continuous voice recording of all answering contents of the person being asked in the complete interactive process from the start to the end of the questioning and answering procedure. The whole answering recording is obtained through the recording of the questioning and answering terminal. The recording software is downloaded in the questioning and answering terminal. The complete transcript refers to the structured and standardized text document formed by combining the questioning contents, follow-up links, key information annotations, etc. in the questioning and answering procedure, and arranging the text after converting the whole answering recording into text through voice recognition technology. The complete transcript corresponding to the whole answering recording can be obtained through the preset complete transcript generation method. The specific complete transcript generation method is described in detail in S530 to S535 below, and is not repeated here. The transcript storage terminal refers to a terminal for storing various complete transcripts. The transcript storage terminal is preset by the person skilled in the art, and is not repeated here.

[0106] Based on the current questioning procedure being consistent with the final questioning procedure, it is indicated that the whole questioning and answering procedure has been completed, and the whole answering recording needs to be collected, the complete transcript is generated, and the complete transcript is uploaded to the transcript storage terminal.

[0107] Referring to Figure 2 , the complete transcript generation method includes the following steps:

[0108] S530: Generating the whole answering content based on the whole answering recording.

[0109] The whole answering content refers to the text information converted from the whole answering recording of the person being asked. The whole answering content can be obtained through voice recognition technology.

[0110] S531: Generating the answering content point in response to the whole answering content.

[0111] The answering content point refers to the smallest independent semantic unit extracted from the text content converted from the whole answering recording. The answering content point is obtained by automatically splitting through natural language processing technology.

[0112] S532: Obtaining the content point filling position according to the answering content point and the preset question list.

[0113] The question list refers to a structured template designed in advance, containing a series of standardized questions. The question list is set in advance by those skilled in the art, which is not described here. The content point filling position refers to the position of the question item in the question list that is semantically matched with the answer content point. Through a semantic matching algorithm (such as keyword mapping, vector similarity calculation), the answer content point can be associated with the question item in the question list, and then the specific position where each content point should be filled in is determined to obtain the content point filling position.

[0114] S533: Fill the answer content point into the question list at the content point filling position to obtain a complete question list.

[0115] The complete question list refers to the structured document formed after filling the answer content point into the corresponding position of the question list. The control question and answer terminal fills the answer content point into the question list at the content point filling position to obtain the complete question list.

[0116] S534: Based on the complete question list to generate visual evidence and judgment results, based on the judgment results to conduct intelligent questioning, and collect judgment answer content.

[0117] The visual evidence refers to the information display method of presenting the structured data in the complete question list in a direct form such as charts, graphs, and relationship maps. By understanding the complete question list, the key data fields therein can be known, and the corresponding visual chart of the key data field can be matched from the preset chart database, and the key data field is placed in the visual chart to obtain the visual evidence. The chart database stores different visual charts corresponding to different key data fields. The chart database is formed by sequentially recording different visual charts corresponding to different key data fields by those skilled in the art, which is not described here.

[0118] The judgment result refers to the conclusion judgment obtained after analyzing the question and answer content. Similarly, by understanding the complete question list, the judgment key data field that has an impact on the judgment can be known, and the corresponding judgment result can be queried from the preset judgment reference table. The table records different judgment results corresponding to different judgment key data fields, and the judgment reference table is recorded by sequentially testing different judgment results corresponding to different judgment key data fields by those skilled in the art, which is not described here.

[0119] The judgment answer content refers to the answer content supplemented by the person being asked through intelligent questioning for the matters that need to be further confirmed in the judgment result. The judgment answer content is converted into text content after being recorded by the question and answer terminal.

[0120] When the visual evidence and the judgment result are generated, the question-answering terminal needs to control the corresponding video recording device to open the personnel to intelligently ask questions (such as whether the judgment result is clear) based on the judgment result, and collect the judgment answer content, so as to facilitate the subsequent steps.

[0121] S535: When the judgment answer content is consistent with the preset reference answer content, the complete record is obtained based on the complete question list, the visual evidence, the judgment result, and the judgment answer content.

[0122] The reference answer content refers to the content that the questioned personnel needs to answer (such as the judgment result has been clear). The reference answer content is set by the person skilled in the art in advance, and is not described here.

[0123] When the judgment answer content is consistent with the reference answer content, it means that all processes in this answer link have been completed, and the complete record needs to be generated. The complete record can be obtained by integrating the documents through the complete question list, the visual evidence, the judgment result, and the judgment answer content, and performing the preset structured arrangement. The structured arrangement is set by the person skilled in the art in advance, and is not described here.

[0124] Reference Figure 3 It also includes the step after collecting the question-answering execution signal of the preset question-answering terminal:

[0125] S10: Responding to the question-answering execution signal to collect the question-answering terminal number.

[0126] The question-answering terminal number refers to the number of the question-answering terminal. In this embodiment, there are multiple question-answering terminals for answering, and each question-answering terminal is provided with a corresponding question-answering terminal number. The question-answering terminal number can be queried through the question-answering terminal. The question-answering terminal records its own corresponding question-answering terminal number.

[0127] When the question-answering terminal sends the question-answering execution signal, the question-answering terminal number needs to be collected to facilitate the subsequent steps.

[0128] S11: Generating a terminal use scenario based on the question-answering terminal number.

[0129] The terminal use scenario refers to the specific question-answering scenario in which the question-answering terminal participates. The terminal use scenario corresponding to the question-answering terminal number can be queried through the preset scenario correspondence table, which records different terminal use scenarios corresponding to different question-answering terminal numbers. The scenario correspondence table is set by the person skilled in the art in advance, and is not described here.

[0130] For example, there are two question and answer terminals A and B, and the terminal use scenarios include indoor and outdoor (field use). A corresponds to indoor, and B corresponds to outdoor (field use). Therefore, when the question and answer terminal number is A, it means that the question and answer is conducted indoors. Similarly, when the question and answer terminal number is B, it means that the question and answer is conducted outdoors (field use).

[0131] S12: generating a recording collection method in response to the terminal use scenario.

[0132] The recording collection method refers to a method for recording the answers of the person being questioned. The specific recording collection method is described in detail in subsequent S130 to S1303 and S131 to S1315, and will not be repeated here.

[0133] The recording database can match the recording collection method corresponding to the terminal use scenario. The database stores different recording collection methods corresponding to different terminal use scenarios. The recording database is formed by sequentially recording different recording collection methods corresponding to different terminal use scenarios by those skilled in the art, which will not be repeated here.

[0134] S13: collecting the answer recording of the person being questioned based on the recording collection method.

[0135] The answer recording of the person being questioned is collected by the recording collection method for subsequent use.

[0136] Referring to Figure 4 , the recording collection method includes the following steps:

[0137] S130: collecting indoor image information when the terminal use scenario is a preset indoor scenario.

[0138] The indoor scenario refers to a scenario in which the question and answer session is conducted indoors (such as an inquiry room). The indoor image information refers to the image in the room where the question and answer session is conducted. The indoor image information is obtained by photographing by a preset camera in the room.

[0139] When the terminal use scenario is an indoor scenario, indoor image information needs to be collected for subsequent steps.

[0140] S1300: performing portrait posture recognition on the indoor image information to obtain the current speaking direction of the person.

[0141] The current speaking direction of the person refers to the direction of the person being questioned when speaking. The current speaking direction of the person can be obtained by performing portrait posture recognition on the indoor image information by computer vision technology. The computer vision technology is well known in the art and will not be repeated here.

[0142] S1301: When the current speaking orientation of the person is inconsistent with the preset person reference speaking orientation, the auxiliary sound collection device number is matched according to the current speaking orientation of the person.

[0143] The person reference speaking orientation refers to the orientation in which the person should be when speaking. The person reference speaking orientation is set by a person skilled in the art in advance, and is not described here. The auxiliary sound collection device number refers to the number of the multi-directional auxiliary sound collection device deployed in advance to solve the problem of poor sound collection caused by the deviation of the speaking orientation of the person being questioned from the reference direction. In the room, an array of auxiliary sound collection devices is pre-set, and each auxiliary sound collection device has a corresponding number and is bound to a specific sound collection direction. The auxiliary sound collection device number corresponding to the current speaking orientation of the person can be matched through the preset auxiliary sound collection database. The database stores different auxiliary sound collection device numbers corresponding to different current speaking orientations of the person, and the auxiliary sound collection database is formed by sequentially recording the different auxiliary sound collection device numbers corresponding to the different current speaking orientations of the person by a person skilled in the art, which is not described here.

[0144] When the current speaking orientation of the person is inconsistent with the person reference speaking orientation, it means that the orientation of the person being questioned has changed, and the auxiliary sound collection device number needs to be matched first.

[0145] S1302: Based on the auxiliary sound collection device number, the corresponding auxiliary sound collection device is controlled to be turned on, and the auxiliary collection recording of the auxiliary sound collection device and the preset main collection recording of the main sound collection device are collected.

[0146] The main sound collection device refers to the device used for main sound collection of the person being questioned (i.e. when the person being questioned does not change the orientation). The auxiliary collection recording refers to the recording of the person being questioned collected by the auxiliary sound collection device corresponding to the auxiliary sound collection device number. The main collection recording refers to the recording of the person being questioned collected by the main sound collection device. The auxiliary collection recording can be obtained by calling the auxiliary sound collection terminal connected to the auxiliary sound collection device. Similarly, the main collection recording can be obtained by calling the main sound collection terminal connected to the main sound collection device. The auxiliary sound collection terminal and the main sound collection terminal are set by a person skilled in the art in advance, and are not described here.

[0147] The auxiliary sound collection device corresponding to the auxiliary sound collection device number is controlled to be turned on, and the auxiliary collection recording of the auxiliary sound collection device and the main collection recording of the main sound collection device are collected.

[0148] S1303: According to the auxiliary collection recording and the main collection recording, the answer recording is obtained.

[0149] The auxiliary collected recording and the main collected recording are audio combined, and the answer recording is obtained by processing the two recordings through a multi-channel audio fusion technology. The multi-channel audio fusion technology is common knowledge in the art, and will not be described here.

[0150] With reference to Figure 5 The verification method of the answer recording includes the following steps:

[0151] S13030: Matching the demand recording device number according to the current speaking direction of the person.

[0152] The demand recording device number refers to the number of a multi-directional recording device deployed in advance to cope with the poor sound collection effect caused by the speaking direction of the person being asked deviating from the reference direction in a question and answer scene. In a room, an array of recording devices is pre-set, and each recording device has a corresponding number and is bound to a specific sound collection direction. The current speaking direction of the person corresponding to the demand recording device number can be matched through a preset auxiliary recording database. The database stores different demand recording device numbers corresponding to different current speaking directions of different persons, and the auxiliary recording database is formed by sequentially recording the different demand recording device numbers corresponding to the different current speaking directions of different persons by a person skilled in the art, which will not be described here.

[0153] S13031: Based on the demand recording device number, the corresponding recording device is controlled to be turned on, and the person's speaking video is collected.

[0154] The person's speaking video refers to the video of the person being asked when speaking. The person's speaking video is obtained by recording through the recording device corresponding to the demand recording device number.

[0155] The corresponding recording device of the demand recording device number is controlled to be turned on, and the person's speaking video is collected for subsequent steps.

[0156] S13032: Mouth shape recognition of the person from the person's speaking video to obtain the person's mouth shape action.

[0157] The person's mouth shape action refers to the action of the person's mouth. The person's mouth shape action is obtained by detecting, tracking and key point analysis of the person's mouth region in the video through computer vision technology, and extracting the lip opening frequency, opening amplitude and pronunciation action sequence.

[0158] S13033: Based on the person's mouth shape action, the auxiliary collected recording and the main collected recording, the answer recording is obtained.

[0159] The audio signals of the auxiliary recording and the main recording are converted into Mel spectrograms, and the video sequence of the mouth movement is extracted as a time sequence feature vector. By using a preset dynamic time warping algorithm, the time axes of the Mel spectrogram feature sequence and the mouth movement time sequence feature are aligned, so that the audio and video time sequences are accurately aligned. Finally, based on the time-aligned audio and mouth movement features, the answer recording is verified and generated from three dimensions of time consistency, content matching degree and intensity correlation. The dynamic time warping algorithm is a common knowledge in the art, and will not be described here.

[0160] Referring to Figure 6 The recording collection method further includes the following steps:

[0161] S131: When the terminal use scenario is a preset outdoor scenario, collect the recording signal of the preset outdoor recording device.

[0162] The outdoor scenario refers to a scenario in which the question and answer session is conducted outdoors. The outdoor recording device refers to a recording device that needs to be used when the question and answer session is conducted outdoors. The recording signal refers to a signal emitted by the outdoor recording device for detecting whether the orientation of the person being questioned changes. The recording signal is emitted by the signal transceiver on the outdoor recording device after the signal trigger button on the outdoor recording device is triggered by the person asking questions.

[0163] When the terminal use scenario is an outdoor scenario, the recording signal of the outdoor recording device needs to be collected for subsequent steps.

[0164] S1310: When the recording signal is consistent with the preset recording adjustment signal, the orientation of the person being questioned is adjusted based on the recording signal.

[0165] The recording adjustment signal refers to a signal when the orientation of the person being questioned changes. The recording adjustment signal is set by a person skilled in the art in advance and will not be described here. The orientation adjustment of the person refers to the adjusted facial orientation of the person being questioned. The orientation adjustment of the person corresponding to the recording signal can be queried through a preset signal-orientation correspondence table, which records different orientation adjustments of the person corresponding to different recording signals. The signal-orientation correspondence table is formed by sequentially recording different orientation adjustments of the person corresponding to different recording signals by a person skilled in the art, and will not be described here. For example, the outdoor recording device can emit three signals A, B and C, wherein A and B are recording adjustment signals. When the signal is A, the signal-orientation correspondence table corresponds to the person being questioned adjusting the orientation to the left. When the signal is B, the signal-orientation correspondence table corresponds to the person being questioned adjusting the orientation to the right.

[0166] When the recording signal is consistent with the recording adjustment signal, it indicates that the orientation of the person being questioned changes, and the orientation adjustment of the person needs to be queried first for subsequent steps.

[0167] S1311: Adjust the orientation of the person to match the specific open window of the vehicle.

[0168] The specific open window of the vehicle refers to the specific open window position of the vehicle when the adjusted facial orientation of the person being questioned is facing or close to it. The specific open window corresponding to the adjusted orientation of the person can be queried through a preset orientation-window correspondence table, which records different specific open windows corresponding to different adjusted orientations of the person. The orientation-window correspondence table is formed by sequentially recording the different specific open windows corresponding to the different adjusted orientations of the person by the person skilled in the art, and will not be described here.

[0169] In this embodiment, the person being questioned will be taken to the side of the vehicle (between the two front and rear doors) for questioning.

[0170] S1312: Based on the specific open window to get the vehicle recording device number.

[0171] The vehicle recording device number refers to the number of recording devices deployed at different positions of the vehicle. The specific open window corresponding to the vehicle recording device number can be queried through a preset window-device correspondence table, which records different vehicle recording device numbers corresponding to different specific open windows. The window-device correspondence table is formed by sequentially recording the different vehicle recording device numbers corresponding to the different specific open windows by the person skilled in the art, and will not be described here.

[0172] S1313: Based on the vehicle recording device number to control the corresponding vehicle recording device to start, and open the window corresponding to the specific open window, while controlling the preset recording masking device to mask the vehicle recording device.

[0173] The vehicle recording device refers to the recording device deployed on the vehicle. The recording masking device refers to the device used to mask the vehicle recording device. The vehicle recording device corresponding to the vehicle recording device number is controlled to start, and the window corresponding to the specific open window is opened, while the recording masking device is controlled to mask the vehicle recording device, in order to assist in sound collection.

[0174] S1314: Collect vehicle collection recordings of the vehicle recording device and outdoor collection recordings of the outdoor recording device.

[0175] The vehicle collection recording refers to the recording of the person being questioned collected by the vehicle recording device. The outdoor collection recording refers to the recording of the person being questioned collected by the outdoor recording device. The vehicle collection recording can be retrieved from the vehicle sound terminal connected to the vehicle recording device. Similarly, the outdoor collection recording can be retrieved from the outdoor sound terminal connected to the outdoor recording device. The vehicle sound terminal and the outdoor sound terminal are set by the person skilled in the art in advance, and will not be described here.

[0176] S1315: Collecting the recording in the vehicle and the recording outdoors to obtain the answer recording.

[0177] This step is the same as S1303 described above, and will not be repeated here.

[0178] Referring to Figure 7 The method for generating the answer content includes the following steps:

[0179] S30: Generating the recording text from the answer recording.

[0180] The recording text refers to the textual representation of the audio information in the answer recording. The recording text corresponding to the answer recording can be obtained through speech recognition technology. Speech recognition technology is well known in the art and will not be repeated here.

[0181] S31: Controlling the pre-set question and answer terminal to perform sentence detection on the recording text to obtain a sentence fluency detection result.

[0182] The sentence fluency detection result refers to the judgment conclusion about whether the text sentence is smooth and the logic is coherent after analyzing the syntax, semantics and expression fluency of the recording text. The sentence fluency detection result can be obtained by controlling the question and answer terminal to perform sentence detection on the recording text using natural language processing technology. Natural language processing technology is well known in the art and will not be repeated here.

[0183] S32: Determining whether the sentence fluency detection result contains a pre-set sentence stuttering feature.

[0184] The sentence stuttering feature refers to a feature identifier for identifying that there is a sentence that is not smooth, interrupted or unnatural pause in the recording text.

[0185] For example, the pause time between words exceeds the average length of normal communication (e.g., more than 2 seconds), the same word or phrase is repeated without meaning (e.g., "I need…"), non-semantic filler words appear frequently (e.g., "um" "er" "that" etc. appear multiple times in a short period of time (e.g., ≥3 times in every 10 words)), the sentence structure is incomplete or the logic is broken (e.g., "Today the weather is very good, then… I want to go…") and the like. The specific sentence stuttering feature is set by a person skilled in the art in advance and will not be repeated here.

[0186] By judging whether the sentence fluency detection result contains the sentence stuttering feature, it can be known whether the recording text can be directly output to obtain the answer content.

[0187] S320: When the sentence fluency detection result contains the sentence stuttering feature, performing stuttering processing using a pre-set stuttering processing method.

[0188] The stutter processing method refers to a method for removing the stutter characteristics of the recorded text. The specific stutter processing method is described in detail in subsequent S3200 to S3205, which will not be repeated here.

[0189] When the stutter characteristics are included in the sentence fluency detection result, it means that the recorded text cannot be directly output to obtain the answer content, and the stutter processing method needs to be used to process the part of the recorded text with stutter characteristics.

[0190] S321: When the stutter characteristics are not included in the sentence fluency detection result, the recorded text is used to obtain the answer content.

[0191] When the stutter characteristics are not included in the sentence fluency detection result, the recorded text can be directly output to obtain the answer content. Since the answer content is also a text content, the recorded text can be directly converted into the answer content.

[0192] Reference Figure 8 The stutter processing method includes the following steps:

[0193] S3200: According to the sentence fluency detection result, the stutter sentence is retrieved.

[0194] The stutter sentence refers to a sentence in the text of the recorded text that has stutter characteristics. The stutter sentence can be retrieved from the sentence fluency detection result. The stutter sentence is included in the sentence fluency detection result.

[0195] S3201: Based on the recorded text and the stutter sentence, the stutter timeline is obtained.

[0196] The stutter timeline refers to the period when the stutter characteristics appear in the recording. The stutter sentence is aligned with the time axis of the answer recording to extract the period when the stutter occurs, and then the stutter timeline is obtained.

[0197] S3202: Collect the answer recording video.

[0198] The answer recording video refers to a video collected synchronously with the answer recording. The answer recording video is recorded by a camera.

[0199] S3203: Based on the answer recording video and the stutter timeline, the stutter segment is determined.

[0200] The stutter segment refers to a video segment corresponding to the stutter timeline extracted from the answer recording video. The stutter segment is obtained by extracting the corresponding stutter timeline from the answer recording video.

[0201] S3204: Mouth recognition is performed on the characters in the stutter segment to obtain the character stutter mouth shape, and based on the character stutter mouth shape and the recorded text, the specific stutter content is obtained.

[0202] The character mouth shape refers to the movement state and morphological characteristics of the mouth of a character in a mouthpiece segment. Through computer vision technology, the mouth shape of a character in a mouthpiece segment can be obtained after detecting, tracking and analyzing the mouth movement of the character. The computer vision technology is well known in the art and will not be described here. The specific mouthpiece content refers to the real semantic content that is not fully expressed or repeated by the person being asked during the voice mouthpiece period. Through a preset mouth shape-phoneme correspondence table, the phoneme sequence corresponding to the character mouth shape can be queried, and then the phoneme sequence is combined with the recorded text to decode the semantic content according to a preset language model, so as to obtain the specific mouthpiece content. The specific language model is set by a person skilled in the art in advance and will not be described here.

[0203] The mouth shape-phoneme correspondence table records different phoneme sequences corresponding to different character mouth shapes. The mouth shape-phoneme correspondence table is formed by recording the different phoneme sequences corresponding to different character mouth shapes after being tested by a person skilled in the art, and will not be described here.

[0204] S3205: obtaining a smooth sentence based on the mouthpiece sentence and the specific mouthpiece content.

[0205] The smooth sentence refers to a text sentence with complete grammar, coherent logic and natural expression. By deleting, merging or completing the specific mouthpiece content (such as filler words, repeated segments and interrupted semantics) in the mouthpiece sentence, and combining with the context to optimize the expression, a smooth sentence can be obtained.

[0206] For example, the repeated content "I I I" and the filler word "um" in the mouthpiece sentence "I I I … um … went to the supermarket" are replaced by the real semantic "I" corresponding to the specific mouthpiece content, and the completed smooth sentence is "I went to the supermarket".

[0207] The method further comprises the following steps:

[0208] S60: collecting an answer recording video.

[0209] This step is the same as S3202 described above and will not be described here.

[0210] S61: performing mouth shape recognition on a character from the answer recording video to obtain a video character mouth shape.

[0211] The video character mouth shape refers to the mouth movement of the person being asked in the answer recording video. Through computer vision technology, the video character mouth shape can be obtained after detecting, tracking and analyzing the mouth movement of the character in the answer recording video.

[0212] S62: obtaining a video text content based on the video character mouth shape.

[0213] The video text content refers to the text content corresponding to the pronunciation of a character inferred by analyzing the dynamic characteristics and morphological characteristics of the mouth shape of the character in the video through a preset lip-reading technology. The lip-reading technology is common knowledge in the art, and will not be described here.

[0214] S63: reporting a forbidden word abnormality prompt when and only when the video text content contains a preset forbidden word.

[0215] The forbidden word refers to a sensitive word or phrase that is prohibited from appearing in the question and answer session. The specific forbidden word is set by a person skilled in the art in advance, and will not be described here.

[0216] When and only when the video text content contains a forbidden word, a forbidden word abnormality prompt needs to be reported to prompt the person asking questions.

[0217] Based on the same inventive concept, the present application provides an AI-based intelligent question and answer scenario generation system, comprising:

[0218] The acquisition module is used to acquire the question and answer execution signal, the answer recording, the current question process, the whole scene answer recording, the judgment answer content, the question and answer terminal number, the indoor image information, the auxiliary collection recording, the main collection recording, the person speaking video, the recording signal, the vehicle collection recording, the outdoor collection recording, and the answer recording video.

[0219] The memory is used to store the program of the AI-based intelligent question and answer scenario generation method.

[0220] The processor is used to load and execute the program stored in the memory.

[0221] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0222] The above is only the preferred embodiment of the present application, and the protection scope of the present application is not limited to the above-mentioned embodiments. Any technical solution falling within the scope of the present application is within the protection scope of the present application. It should be noted that, for those skilled in the art, some improvements and refinements without departing from the principles of the present application are also considered within the protection scope of the present application.

Claims

1. An AI-based intelligent question and answer scenario generation method, characterized by, The method comprises the following steps: S1: collecting a preset question and answer terminal question and answer execution signal; S2: collecting the respondent's answer recording in response to the question and answer execution signal; S3: generating answer content and content keywords based on the answer recording; S4: matching the question and answer scene based on the content keywords; S5: generating a reference answer content and a final questioning process in response to the question and answer scene; S50: when the answer content is inconsistent with the reference answer content, obtaining the answer missing content according to the answer content and the reference answer content; S51: generating a follow-up question based on the answer missing content, and conducting follow-up questioning based on the follow-up question; S52: when the answer content is consistent with the reference answer content or after the follow-up questioning is completed, reporting the questioning process for further prompting and collecting the current questioning process; S53: when the current questioning process is consistent with the final questioning process, collecting the entire scene answer recording, generating a complete transcript based on the entire scene answer recording, and uploading the complete transcript to a preset transcript storage terminal; Further comprising the steps after collecting the question and answer execution signal of the preset question and answer terminal: S10: collecting the question and answer terminal number in response to the question and answer execution signal; S11: generating a terminal use scene based on the question and answer terminal number; S12: generating a recording collection method in response to the terminal use scene; S13: collecting the respondent's answer recording based on the recording collection method; The recording collection method further comprises: S131: when the terminal use scene is a preset outdoor scene, collecting a preset outdoor recording device recording signal; S1310: when the recording signal is consistent with a preset recording adjustment signal, adjusting the orientation of the personnel based on the recording signal; S1311: matching the specific opening of the car window according to the personnel adjustment orientation; S1312: obtaining the vehicle recording device number based on the specific opening of the car window; S1313: controlling the corresponding vehicle recording device to be turned on based on the vehicle recording device number, opening the car window corresponding to the specific opening of the car window, and controlling the preset recording shielding device to shield the vehicle recording device; S1314: collecting the vehicle collection recording of the vehicle recording device and the outdoor collection recording of the outdoor recording device; S1315: obtaining the answer recording according to the vehicle collection recording and the outdoor collection recording. 2.The AI-based intelligent question and answer scenario generation method of claim 1, wherein, Further comprising a complete transcript generation method: S530: generating an entire scene answer content based on the entire scene answer recording; S531: generating an answer content point in response to the entire scene answer content; S532: obtaining a content point filling position according to the answer content point and a preset question list; S533: filling the answer content point into the content point filling position in the question list to obtain a complete question list; S534: generating a visual evidence and a judgment result based on the complete question list, conducting intelligent questioning based on the judgment result, and collecting a judgment answer content; S535: when the decision answer content is consistent with the preset reference answer content, obtaining complete notes based on the complete question list, the visual evidence, the decision result, and the decision answer content. 3.The AI-based intelligent question and answer scenario generation method of claim 1, wherein, The audio collection method comprises: S130: when the terminal use scene is a preset indoor scene, collecting indoor image information; S1300: performing portrait posture recognition on the indoor image information to obtain a current speaking orientation of a person; S1301: when the current speaking orientation of the person is inconsistent with a preset reference speaking orientation of a person, matching an auxiliary sound collection device number according to the current speaking orientation of the person; S1302: based on the auxiliary sound collection device number, controlling the corresponding auxiliary sound collection device to be turned on, collecting auxiliary collection audio of the auxiliary sound collection device, and collecting main collection audio of a preset main sound collection device; S1303: obtaining answer audio according to the auxiliary collection audio and the main collection audio.

4. The AI-based intelligent question and answer scenario generation method of claim 3, wherein, Also included is a verification method for the answer audio: S13030: matching a demand video recording device number according to the current speaking orientation of the person; S13031: based on the demand video recording device number, controlling the corresponding video recording device to be turned on, and collecting a speaking video of the person; S13032: performing lip recognition on the person from the speaking video of the person to obtain a lip action of the person; S13033: obtaining answer audio based on the lip action of the person, the auxiliary collection audio, and the main collection audio. 5.The AI-based intelligent question and answer scenario generation method of claim 1, wherein, Also included is a generation method for the answer content: S30: generating audio text according to the answer audio; S31: controlling a preset question and answer terminal to perform sentence detection on the audio text to obtain a sentence fluency detection result; S32: determining whether the sentence fluency detection result contains a preset sentence stuttering feature; S320: when the sentence fluency detection result contains the sentence stuttering feature, performing stuttering processing by a preset stuttering processing method; S321: when the sentence fluency detection result does not contain the sentence stuttering feature, obtaining answer content based on the audio text.

6. The AI-based intelligent question and answer scenario generation method of claim 5, wherein, The stuttering processing method comprises: S3200: retrieving a stuttering sentence according to the sentence fluency detection result; S3201: obtaining a stuttering timeline based on the audio text and the stuttering sentence; S3202: collecting answer audio video; S3203: determining a stuttering segment based on the answer audio video and the stuttering timeline; S3204: performing lip recognition on the person from the stuttering segment to obtain a person's stuttering lip shape, and obtaining specific stuttering content based on the person's stuttering lip shape and the audio text; S3205: obtaining a fluent sentence based on the stuttering sentence and the specific stuttering content. 7.The AI-based intelligent question answering scenario generation method of claim 1, wherein, Also included are: S60: collecting answer audio video; S61: performing lip recognition on the person from the answer audio video to obtain a video person's lip shape; S62: obtaining video text content based on the video person's lip shape; S63: reporting a prohibited word abnormality prompt only when the video text content contains a preset prohibited word. 8.An AI-based intelligent question-answering scenario generation system, characterized by, Included are: The collection module is used for collecting question and answer execution signals, answer audios, current question and answer process and whole game answer audios. The memory is used for storing a program of the AI-based intelligent question and answer scene generation method according to any one of claims 1 to 7. The processor is used for loading and executing the program stored in the memory.

Citation Information

Patent Citations

  • Guided question and answer alarm receiving and processing method based on NLP speech recognition and AI agent technology

    CN120017751A