A conference assistance system based on information analysis

By using the mouth feature extraction module and the audio analysis module in the conference auxiliary system, the problem of the inability to accurately identify the user's speech time period in the prior art is solved, and the precise extraction and efficient processing of the conference content are achieved.

CN119831555BActive Publication Date: 2025-06-06ZHONGHE YUNKE INFORMATION TECH GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510299813.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-06
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The prior art cannot accurately identify the user's speech time period, resulting in the inability to accurately extract the user's speech content, affecting the efficiency of the extraction of conference content.

Method used

The information acquisition module synchronously obtains the video and audio information of each user. The feature extraction module extracts the mouth features based on the video information and issues interception analysis instructions for the audio information. The audio extraction module intercepts and converts the audio information. The calibration division module calibrates the interception analysis instructions based on the verification and volume integral parameters, analyzes the module periodically counts the number of calibration instructions, and adjusts the parameters obtained by the information to ensure the content integrity.

Benefits of technology

It realizes accurate identification of user speech time periods and accurate extraction of content, improving the efficiency of conference content extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119831555B_ABST
    Figure CN119831555B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of conference management, and in particular to a conference auxiliary system based on information analysis. The system comprises an information collection module, a feature extraction module, an audio extraction module, a calibration division module and an analysis module. The system analyzes audio information corresponding to a single interception and analysis instruction to determine a verification parameter and a volume integral parameter, and calibrates each interception and analysis instruction based on the verification parameter and the volume integral parameter; periodically determines whether the acquisition of information is qualified based on the number of each interception and analysis instruction calibrated by statistics; when it is determined that the acquisition of information is abnormal, timely adjusts the parameters of the information acquisition to completely acquire the text information of the meeting, accurately extract the user's speech content, and improve the efficiency of extracting the meeting content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of conference management, and in particular to a conference assistance system based on information analysis. Background Art

[0002] Meetings are an important form of communication and decision-making in organizations. With the development of information technology, meetings have become characterized by large amounts of information and professional content, which has put forward new requirements for the analysis and management of meeting content. Effectively extracting key meeting content, identifying professional terms, and combining audio-visual information for comprehensive presentation can help participants fully understand the key points of the meeting and improve meeting efficiency.

[0003] The existing meeting record method mainly relies on manual work, such as manual recording of minutes by secretaries or participants, and video recording. Audio and video recording can retain the original data, but content extraction requires manual review, which is labor-intensive and inefficient. This situation needs to be further improved.

[0004] Chinese Patent Publication No.: CN118660128A, discloses a conference content analysis method, system and conference all-in-one machine based on artificial intelligence, including collecting video and audio data of the conference site, using video recognition technology to extract presentation information from the video, and then based on the extracted presentation information, using a pre-trained domain classification model to determine the professional field to which the conference belongs, according to the determined field, calling the corresponding domain language model to perform speech recognition on the audio, converting the recognition result into a conference speech text, and using domain knowledge to mark the proprietary terms in the text, and finally, extracting the key content of the conference speech text and presenting it on the conference all-in-one machine; it can be seen that the above technical solution has the following problems: it is impossible to accurately identify the user's speech time period, resulting in the inability to accurately extract the user's speech content, thereby affecting the efficiency of extracting the conference content. Summary of the invention

[0005] To this end, the present invention provides a conference assistance system based on information analysis to overcome the problem in the prior art that the user's speech time period cannot be accurately identified, resulting in the inability to accurately extract the user's speech content, thereby affecting the efficiency of extracting the conference content.

[0006] To achieve the above object, the present invention provides a conference assistance system based on information analysis, comprising:

[0007] An information collection module, which is used to store the video information and audio information of each user synchronously acquired;

[0008] A feature extraction module, which is connected to the information acquisition module, is used to extract the mouth features of each user based on the acquired video information, and issue an interception and analysis instruction for the audio information according to the mouth aspect ratio of the user's mouth features;

[0009] An audio extraction module, which is connected to the feature extraction module and the information collection module, and is used to intercept audio information based on the interception analysis instruction, and convert the intercepted audio information into text to obtain the conference text content;

[0010] a calibration division module, which is connected to the audio extraction module and the information acquisition module, and is used to determine a verification parameter and a volume integral parameter for a single interception and analysis instruction based on the audio information, and calibrate each interception and analysis instruction based on the verification parameter and the volume integral parameter;

[0011] An analysis module is connected to the calibration division module, the information acquisition module, the feature extraction module and the audio extraction module, and is used to periodically determine whether the acquisition of information is qualified based on the number of each interception and analysis instruction calibrated by statistics. When it is determined that the acquisition of information is abnormal, the instruction issuance standard for issuing the interception and analysis instruction is adjusted to a corresponding value, the bit rate of the audio information transmitted by the information acquisition module is adjusted to a corresponding value, or the time node corresponding to the starting analysis instruction of the interception and analysis instruction is preset in advance for the acquisition duration.

[0012] Furthermore, the interception analysis instruction includes a start analysis instruction and a stop analysis instruction, and the feature extraction module issues a start analysis instruction for the audio information under the condition that the mouth aspect ratio of the user's mouth feature is greater than a preset deformation parameter;

[0013] Obtaining the duration when the aspect ratio of the mouth is less than or equal to the preset deformation parameter, and issuing a termination analysis instruction for the audio information when the duration is greater than the preset holding time;

[0014] Determine the duration between the time point when the mouth aspect ratio of the user's mouth feature is greater than the preset deformation parameter and the time point when the termination analysis instruction is issued for the audio information as the feature change duration;

[0015] For a single interception analysis instruction, the audio extraction module intercepts the audio information within the corresponding time period based on the interception analysis instruction, and determines the duration of the audio within the corresponding time period;

[0016] The calibration division module is used to calculate the ratio of the feature change duration to the audio duration, and record it as a verification parameter;

[0017] The calibration division module is used to calibrate the interception analysis instruction based on the verification parameter, including:

[0018] If the verification parameter is less than or equal to the first preset verification parameter, the intercepted analysis instruction is marked as a qualified instruction;

[0019] If the verification parameter is less than or equal to the second preset verification parameter and greater than the first preset verification parameter, the interception analysis instruction is calibrated based on the interception duration of the corresponding time period intercepted by the interception analysis instruction;

[0020] If the verification parameter is greater than the second preset verification parameter, the interception analysis instruction is calibrated based on the volume integral parameter.

[0021] Furthermore, the calibration division module is used to calibrate the interception analysis instructions based on the interception duration, including:

[0022] If the interception time is less than or equal to the preset interception time, the interception analysis instruction is calibrated based on the volume integral parameter;

[0023] If the interception duration is longer than the preset interception duration, the interception analysis instruction will be marked as a low-sensitivity instruction.

[0024] Furthermore, the calibration division module is used to calibrate the interception analysis instruction based on the volume integral parameter, including:

[0025] Determine a volume integral parameter, draw a volume time domain curve based on the volume of the audio information corresponding to the single interception analysis instruction, and record the calculated integral value of the volume time domain curve as the volume integral parameter;

[0026] If the volume integral parameter is less than or equal to the preset volume integral parameter, the interception analysis instruction is marked as a fluctuation instruction;

[0027] If the volume integral parameter is greater than the preset volume integral parameter, the intercepted analysis instruction is marked as a delay instruction.

[0028] Furthermore, the analysis module is used to determine whether the acquisition of information is qualified based on the statistical number of the calibrated interception analysis instructions, including:

[0029] If the number of qualified instructions is the maximum number in a single cycle, it is determined that the acquisition of information is qualified, and the video information and audio information are continuously analyzed;

[0030] If the number of low-sensitivity instructions is the maximum number in a single cycle, it is determined that the acquisition of information is abnormal, and the first preset check parameter and the second preset check parameter are adjusted based on the variance of the interception time length of each low-sensitivity instruction;

[0031] If the number of fluctuation instructions is the maximum number in a single cycle, it is determined that the acquisition of information is abnormal, and the processing parameters of each information acquired by the information acquisition module are adjusted based on the audio interruption parameter;

[0032] If the number of delayed instructions is the maximum number in a single cycle, it is determined that the acquisition of information is abnormal, and the time node corresponding to the start analysis instruction is preset in advance for the acquisition time based on the average value of the verification parameters of each delayed instruction;

[0033] The interception analysis instruction includes a start analysis instruction and a stop analysis instruction for obtaining a corresponding interception time period.

[0034] Further, the analysis module is used to adjust the first preset check parameter and the second preset check parameter based on the variance of each interception time length of each low-sensitivity instruction, including:

[0035] If the variance is less than or equal to the preset variance, adjusting the first preset check parameter and the second preset check parameter to corresponding values ​​based on the average value of each interception duration of each low-sensitivity instruction;

[0036] If the variance is greater than the preset variance, the first preset verification parameter and the second preset verification parameter of the user corresponding to each low-sensitivity instruction are adjusted to corresponding values ​​based on the average value of each interception time length of each low-sensitivity instruction.

[0037] Further, the analysis module is used to adjust the first preset check parameter and the second preset check parameter to corresponding values ​​based on the average value of each interception time length of each low-sensitivity instruction;

[0038] The reduction range of the first preset check parameter and the second preset check parameter is proportional to the average value of each interception time length of each low-sensitivity instruction.

[0039] Furthermore, the analysis module is used to adjust the processing parameters of each information acquired by the information acquisition module based on the audio interruption parameter, including:

[0040] Obtain the number of intersections between the volume time domain curve of a single fluctuation instruction and the time axis, and record the average value of the calculated number of intersections of each fluctuation instruction as the audio interruption parameter;

[0041] If the audio interruption parameter is greater than the preset interruption parameter, then based on the audio interruption parameter, the instruction issuance standard for issuing the interception analysis instruction for the audio information according to the mouth aspect ratio of the user's mouth features is adjusted to a corresponding value;

[0042] If the audio interruption parameter is less than or equal to the preset interruption parameter, the bit rate of the audio information transmitted by the information acquisition module is adjusted to a corresponding value based on the average value of the volume integral parameters of each fluctuation instruction.

[0043] Furthermore, the analysis module is used to adjust the instruction issuance standard for issuing the interception analysis instruction for the audio information according to the mouth aspect ratio of the user's mouth feature to a corresponding value based on the audio interruption parameter, wherein:

[0044] The increase in the command issuance standard is proportional to the audio interruption parameter.

[0045] Furthermore, the analysis module is used to adjust the bit rate of the audio information transmitted by the information acquisition module to a corresponding value based on the average value of the volume integral parameter of each fluctuation instruction, wherein:

[0046] The increase in the bit rate is inversely proportional to the average value of the volume integral parameter of each fluctuation instruction;

[0047] Based on the average value of the verification parameters of each delay instruction, the time node corresponding to the start analysis instruction is preset in advance for the acquisition time, wherein,

[0048] The increase of the determined preset acquisition time length is proportional to the average value of the verification parameters of each delay instruction.

[0049] Compared with the prior art, the beneficial effect of the present invention lies in that the audio information corresponding to a single interception and analysis instruction is analyzed to determine the verification parameter and the volume integral parameter, and each interception and analysis instruction is calibrated based on the verification parameter and the volume integral parameter; the number of each interception and analysis instruction calibrated based on statistics is periodically determined to determine whether the acquisition of information is qualified; when it is determined that the acquisition of information is abnormal, the parameters of information acquisition are adjusted in time to completely obtain the text information of the meeting, accurately identify the user's speaking time period through mouth features, and accurately extract the user's speech content, thereby improving the efficiency of extracting the meeting content.

[0050] Furthermore, based on the aspect ratio of the user's mouth features, an interception and analysis instruction for the audio information is issued, and the corresponding audio information is intercepted and converted to text analysis only when the user's speaking signal is received, which effectively improves the data conversion and processing efficiency. The interception and analysis instruction is calibrated based on the verification parameter, and the verification parameter characterizes the completeness of the voice content of the corresponding user recorded in the intercepted time period corresponding to the interception and analysis instruction. When the verification parameter is less than or equal to the second preset verification parameter and greater than the first preset verification parameter, there is partial missing of the voice content in the intercepted time period corresponding to the interception and analysis instruction. In this case, the interception and analysis instruction is calibrated in combination with the interception time length. When the interception time length is greater than the preset interception time length, the user's speaking time is relatively long. In this case, the evaluation standard for calibrating the interception and analysis instruction based on the verification parameter should be improved to ensure the complete acquisition of key content. The instruction in this case is calibrated as a low-sensitivity instruction, which further improves the efficiency of extracting meeting content.

[0051] Furthermore, based on the calibration of the volume integral parameter, the interception and analysis instructions are characterized by the volume of the audio information obtained by the audio extraction module for conversion into text information. When the volume integral parameter is less than or equal to the preset volume integral parameter, the audio information fluctuates greatly in this case, resulting in the inability to accurately convert the user's audio information into text. The interception and analysis instructions in this case are calibrated as fluctuation instructions; when the volume integral parameter is greater than the preset volume integral parameter, the audio recording and transmission are normal at this time. At this time, the intercepted audio information is incomplete due to the delay in the interception and analysis instructions. The interception and analysis instructions in this case are calibrated as delay instructions. The instructions in each case are calibrated as a basis for evaluating whether the acquisition of subsequent analysis information is qualified. While accurately determining the specific circumstances of the acquisition of each information, the efficiency of extracting the meeting content is further improved.

[0052] Furthermore, based on the statistically calibrated number of intercepted and analyzed instructions, it is determined whether the acquisition of information is qualified. When the number of low-sensitivity instructions is the maximum number in a single cycle, the variance of each intercepted time is determined. The variance characterizes the fluctuation of the speaking time of each user. When the variance is greater than the preset variance, there are significant differences in the historical speaking habits of each user. In this case, the first preset verification parameter and the second preset verification parameter of the corresponding user are adjusted to improve the sensitivity of instruction calibration, thereby further improving the analysis efficiency of the instruction. When the variance is less than or equal to the preset variance, the historical speaking time of each user is generally similar and is in the discussion stage. At this time, the first preset verification parameter and the second preset verification parameter of each user are adjusted, and the analysis parameters for the instruction are adjusted according to the acquired data. While ensuring efficient and accurate division of the specific circumstances of the instruction, the efficiency of extracting the meeting content is further improved.

[0053] Furthermore, when the number of fluctuation instructions is the maximum number within a single cycle, the processing parameters of each information obtained by the information collection module are adjusted based on the audio interruption parameter. The audio interruption parameter characterizes the interruption of the audio information. When the audio interruption parameter is greater than the preset interruption parameter, the audio information has a high-frequency interruption. At this time, due to the abnormality of the instruction issuance standard for intercepting and analyzing instructions, multiple audio information is still intercepted when the user is speaking. At this time, the instruction issuance standard for the analysis instruction is adjusted to improve the accuracy of effective audio information interception and acquisition; when the audio interruption parameter is less than or equal to the preset interruption parameter, the information has a low-frequency interruption. At this time, the information is lost due to transmission abnormalities. At this time, by adjusting the bit rate parameter, the currently used bit rate value is increased, the audio data transmission speed is increased, and the sound interruption caused by data loss is reduced. While fully acquiring the text information of the meeting, the user's speech content is further accurately extracted.

[0054] Furthermore, when the number of delay instructions is the maximum number within a single cycle, in this case, the intercepted audio information is incomplete due to the presence of a large number of interception and analysis instruction delays. In this case, the acquisition time of the time node corresponding to the start analysis instruction is preset in advance, and the preset acquisition time is further determined based on the verification parameters of each delay instruction. While accurately determining the specific circumstances of each information acquisition, the efficiency of extracting the conference content is further improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 A module block diagram of a conference assistance system based on information analysis according to an embodiment of the present invention;

[0056] Figure 2 A logic decision diagram for calibrating the interception and analysis instructions based on the verification parameters by the calibration and division module of the embodiment of the present invention;

[0057] Figure 3 A logic decision diagram for calibrating and analyzing instructions based on the interception duration by the calibration and division module in an embodiment of the present invention;

[0058] Figure 4 It is a logic decision diagram of the calibration and division module of the embodiment of the present invention based on the volume integral parameter calibration interception analysis instruction. DETAILED DESCRIPTION

[0059] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0060] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0061] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is merely for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0062] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0063] See also Figure 1 , Figure 2 , Figure 3 as well as Figure 4 As shown, they are respectively a module block diagram of a conference auxiliary system based on information analysis according to an embodiment of the present invention, a logic decision diagram of a calibration division module calibrating an interception analysis instruction based on a verification parameter, a logic decision diagram of a calibration division module calibrating an interception analysis instruction based on an interception duration, and a logic decision diagram of a calibration division module calibrating an interception analysis instruction based on a volume integral parameter; a conference auxiliary system based on information analysis according to an embodiment of the present invention comprises:

[0064] An information collection module, which is used to store the video information and audio information of each user synchronously acquired;

[0065] A feature extraction module, which is connected to the information acquisition module, is used to extract the mouth features of each user based on the acquired video information, and issue an interception and analysis instruction for the audio information according to the mouth aspect ratio of the user's mouth features;

[0066] An audio extraction module, which is connected to the feature extraction module and the information collection module, and is used to intercept the audio information within the corresponding time period when receiving the interception analysis instruction, and convert the intercepted audio information into text to obtain the conference text content;

[0067] a calibration division module, which is connected to the audio extraction module and the information acquisition module, and is used to determine a verification parameter and a volume integral parameter for a single interception and analysis instruction based on the corresponding audio information, and calibrate each interception and analysis instruction based on the verification parameter and the volume integral parameter;

[0068] The analysis module is connected to the calibration division module, the information acquisition module, the feature extraction module and the audio extraction module, and is used to periodically determine whether the acquisition of information is qualified based on the number of each interception and analysis instruction of the statistical calibration. When it is determined that the acquisition of information is abnormal, the instruction issuance standard for issuing interception and analysis instructions for audio information based on the mouth aspect ratio of the user's mouth features is adjusted to a corresponding value, the bit rate of the audio information transmitted by the information acquisition module is adjusted to a corresponding value, or the time node corresponding to the starting analysis instruction of the interception and analysis instruction is preset in advance for the acquisition duration.

[0069] Specifically, the specific structure of the information collection module is not limited. It can be a number of cameras and microphones respectively set in various directions in the conference room to obtain video information and audio information of each user. This is existing technology and will not be described in detail.

[0070] Specifically, the information acquisition module is also used to process and convert audio information, including audio encoding output. The information acquisition module includes an audio encoder component, which is actually embodied as an encoder software library. This audio encoder component is used to encode the original audio information according to a specific encoding format and parameters to store and transmit the audio information.

[0071] Specifically, there is no limitation on the specific method of extracting the mouth features of each user. Dlib or OpenCV library can be used to detect facial key points. To extract the key points of the mouth, Dlib's shape_predictor_68_face_landmarks.dat model can be used to detect 68 key points of the face, where the key points of the mouth are 48 to 67. The mouth aspect ratio MAR is determined by calculating the distance between these key points. The mouth aspect ratio is the ratio of the height to the width of the mouth. The distance between the 51st and 59th points in the Dlib face key point detection can be used as the mouth height, and the distance between the 49th and 55th points in Dlib can be used as the mouth width. This will not be repeated.

[0072] Specifically, the interception analysis instruction includes a start analysis instruction and a stop analysis instruction; the feature extraction module issues a start analysis instruction for the audio information under the condition that the mouth aspect ratio of the user's mouth feature is greater than a preset deformation parameter; obtains the duration when the mouth aspect ratio is less than or equal to the preset deformation parameter, and issues a stop analysis instruction for the audio information when the duration is greater than the preset holding time;

[0073] The duration from the time point when the mouth aspect ratio of the user's mouth feature is greater than the preset deformation parameter to the time point when the instruction to terminate the analysis of the audio information is issued is the feature change duration;

[0074] Specifically, the preset deformation parameter is selected within the interval [0.5, 0.7].

[0075] Specifically, the audio information corresponding to a single interception and analysis instruction is analyzed to determine the verification parameters and the volume integral parameters, and each interception and analysis instruction is calibrated based on the verification parameters and the volume integral parameters; the number of each interception and analysis instruction calibrated based on statistics is periodically used to determine whether the acquisition of information is qualified; when it is determined that the acquisition of information is abnormal, the parameters for information acquisition are adjusted in time to completely obtain the text information of the meeting, accurately extract the user's speech content, and improve the efficiency of extracting the meeting content.

[0076] Specifically, the analysis module also includes a statistical unit, a voiceprint recognition unit and a report generation unit;

[0077] The statistical unit pre-stores: the scheduled start time of the meeting and the voiceprint information of each speaker;

[0078] The statistical unit determines the time point at which the start of receiving the audio information is determined as the actual start time;

[0079] The voiceprint recognition unit is used to process the audio information using speech recognition technology, convert the audio into text content, identify the speech content of different voiceprint information, and then match the voiceprint information corresponding to each speaker pre-stored in the statistical unit by analyzing the voice features, so as to identify each speaker;

[0080] For each identified speaker, the statistical unit records the time of his / her first speech; the first speech time is determined as the time point when the corresponding participant starts to participate in the meeting;

[0081] Determine whether each participant is late based on the time when the participant starts to participate in the meeting and the scheduled start time of the meeting; compare the first speaking time of each speaker with the scheduled start time of the meeting. If the first speaking time of a speaker is later than the scheduled start time of the meeting and exceeds the predetermined delay range, the participant is determined to be late.

[0082] The report generation unit is used to compile the list of late participants determined by the voiceprint recognition unit into a report, recording the name of each latecomer and the length of time of lateness, so as to facilitate the conference organizer to view and handle.

[0083] Specifically, for a single interception analysis instruction, the audio extraction module is used to intercept the audio information within the corresponding time period when receiving the interception analysis instruction, and determine the duration of the audio within the corresponding time period;

[0084] The calibration division module is used to calculate the ratio of the feature change duration to the audio duration, and record it as a verification parameter;

[0085] The calibration division module is used to calibrate the interception analysis instruction based on the verification parameter, including:

[0086] If the verification parameter is less than or equal to the first preset verification parameter, the intercepted analysis instruction is marked as a qualified instruction;

[0087] If the verification parameter is less than or equal to the second preset verification parameter and greater than the first preset verification parameter, the interception analysis instruction is calibrated based on the interception duration of the corresponding time period intercepted by the interception analysis instruction;

[0088] If the verification parameter is greater than the second preset verification parameter, the interception analysis instruction is calibrated based on the volume integral parameter.

[0089] Specifically, the first preset calibration parameter A1 is selected within the interval [1, 1.1], and the second preset calibration parameter A2 is selected within the interval [1.35, 1.5].

[0090] Specifically, the audio duration is the sum of the durations of the audios whose volume is not zero during the intercepted time period.

[0091] Specifically, the calibration division module is used to calibrate the interception analysis instructions based on the interception duration, including:

[0092] If the interception time is less than or equal to the preset interception time, the interception analysis instruction is calibrated based on the volume integral parameter;

[0093] If the interception duration is longer than the preset interception duration, the interception analysis instruction will be marked as a low-sensitivity instruction.

[0094] Specifically, the preset interception time length Q0 is selected within the interval [2min, 3min].

[0095] Specifically, based on the aspect ratio of the user's mouth features, an interception and analysis instruction for the audio information is issued, and the corresponding audio information is intercepted only when the user's speaking signal is received for text analysis, which effectively improves the data conversion and processing efficiency, and the interception and analysis instruction is calibrated based on the verification parameter. The verification parameter characterizes the completeness of the voice content of the corresponding user recorded in the intercepted time period corresponding to the interception and analysis instruction. When the verification parameter is less than or equal to the second preset verification parameter and greater than the first preset verification parameter, the voice content in the intercepted time period corresponding to the interception and analysis instruction is partially missing. In this case, the interception and analysis instruction is calibrated in combination with the interception time length. When the interception time length is greater than the preset interception time length, the user's speaking time is relatively long. In this case, the evaluation standard for calibrating the interception and analysis instruction based on the verification parameter should be improved to ensure the complete acquisition of key content, and the instruction in this case is calibrated as a low-sensitivity instruction;

[0096] Specifically, the calibration division module is used to calibrate the interception and analysis instructions based on the volume integral parameter, including:

[0097] Determine a volume integral parameter, draw a volume time domain curve based on the volume of the audio information corresponding to the single interception analysis instruction, and record the calculated integral value of the volume time domain curve as the volume integral parameter;

[0098] If the volume integral parameter is less than or equal to the preset volume integral parameter, the interception analysis instruction is marked as a fluctuation instruction;

[0099] If the volume integral parameter is greater than the preset volume integral parameter, the intercepted analysis instruction is marked as a delay instruction.

[0100] Specifically, the preset volume integral parameter Y0 is selected within the interval [0.46L0, 0.62L0], and L0 is the average value of the volume integral parameters of each historical interception analysis instruction.

[0101] Specifically, the interception and analysis instructions are calibrated based on the volume integral parameter. The volume integral parameter represents the volume of the audio information obtained by the audio extraction module for conversion into text information. When the volume integral parameter is less than or equal to the preset volume integral parameter, the audio information fluctuates greatly, resulting in the inability to accurately convert the user's audio information into text. The interception and analysis instructions in this case are calibrated as fluctuation instructions; when the volume integral parameter is greater than the preset volume integral parameter, the audio recording and transmission are normal at this time. At this time, the intercepted audio information is incomplete due to the delay in the interception and analysis instructions. The interception and analysis instructions in this case are calibrated as delay instructions. The instructions in each case are calibrated as a basis for evaluating whether the acquisition of subsequent analysis information is qualified. While accurately determining the specific circumstances of the acquisition of each information, the efficiency of extracting the meeting content is further improved.

[0102] Specifically, the analysis module is used to determine whether the acquisition of information is qualified based on the number of each interception analysis instruction that is statistically calibrated, including:

[0103] If the number of qualified instructions is the maximum number in a single cycle, it is determined that the acquisition of information is qualified, and the video information and audio information are continuously analyzed;

[0104] If the number of low-sensitivity instructions is the maximum number in a single cycle, it is determined that the acquisition of information is abnormal, and the first preset check parameter and the second preset check parameter are adjusted based on the variance of the interception time length of each low-sensitivity instruction;

[0105] If the number of fluctuation instructions is the maximum number in a single cycle, it is determined that the acquisition of information is abnormal, and the processing parameters of each information acquired by the information acquisition module are adjusted based on the audio interruption parameter;

[0106] If the number of delayed instructions is the maximum number in a single cycle, it is determined that the acquisition of information is abnormal, and the time node corresponding to the start analysis instruction is preset in advance for the acquisition time based on the average value of the verification parameters of each delayed instruction;

[0107] The interception analysis instruction includes a start analysis instruction and a stop analysis instruction for obtaining a corresponding interception time period.

[0108] Specifically, the analysis module is used to adjust the first preset check parameter and the second preset check parameter based on the variance of each interception time length of each low-sensitivity instruction, including:

[0109] If the variance is less than or equal to the preset variance, adjusting the first preset check parameter and the second preset check parameter to corresponding values ​​based on the average value of each interception duration of each low-sensitivity instruction;

[0110] If the variance is greater than the preset variance, the first preset verification parameter and the second preset verification parameter of the user corresponding to each low-sensitivity instruction are adjusted to corresponding values ​​based on the average value of each interception time length of each low-sensitivity instruction.

[0111] Specifically, the preset variance is selected within the interval [0.3T0, 0.47T0], and T0 is the average value of each interception duration of each low-sensitivity instruction.

[0112] Specifically, the number of intercepted and analyzed instructions based on statistical calibration is used to determine whether the acquisition of information is qualified. When the number of low-sensitivity instructions is the maximum number in a single cycle, the variance of each intercepted time is determined. The variance characterizes the fluctuation of the speaking time of each user. When the variance is greater than the preset variance, there are significant differences in the historical speaking habits of each user. In this case, the first preset verification parameter and the second preset verification parameter of the corresponding user are adjusted to improve the sensitivity of instruction calibration, thereby further improving the analysis efficiency of the instruction. When the variance is less than or equal to the preset variance, the historical speaking time of each user is generally similar and is in the discussion stage. At this time, the first preset verification parameter and the second preset verification parameter of each user are adjusted, and the analysis parameters for the instruction are adjusted according to the acquired data. While ensuring efficient and accurate division of the specific circumstances of the instruction, the efficiency of extracting the meeting content is further improved.

[0113] Specifically, the analysis module is used to adjust the first preset check parameter and the second preset check parameter to corresponding values ​​based on the average value of each interception duration of each low-sensitivity instruction;

[0114] The reduction range of the first preset check parameter and the second preset check parameter is proportional to the average value of each interception time length of each low-sensitivity instruction.

[0115] In this embodiment, optionally,

[0116] Recording the average of each intercepted duration as the average duration, and comparing the average duration with the first preset average duration and the second preset average duration;

[0117] If the average duration is less than or equal to the first preset average duration, the first preset calibration parameter is adjusted to 0.92 times the initial first preset calibration parameter, and the second preset calibration parameter is adjusted to 0.91 times the initial second preset calibration parameter;

[0118] If the average duration is less than or equal to the second preset average duration and greater than the first preset average duration, the first preset calibration parameter is adjusted to 0.85 times the initial first preset calibration parameter, and the second preset calibration parameter is adjusted to 0.85 times the initial second preset calibration parameter;

[0119] If the average duration is greater than the second preset average duration, the first preset calibration parameter is adjusted to 0.73 times the initial first preset calibration parameter, and the second preset calibration parameter is adjusted to 0.74 times the initial second preset calibration parameter;

[0120] The first preset average duration is 1.3Q0, and the second preset average duration is 1.7Q0.

[0121] Specifically, the processing parameters of each information acquired by the information acquisition module are adjusted based on the audio interruption parameter, including:

[0122] Obtain the number of intersections between the volume time domain curve of a single fluctuation instruction and the time axis, and record the average value of the calculated number of intersections of each fluctuation instruction as the audio interruption parameter;

[0123] If the audio interruption parameter is greater than the preset interruption parameter, then based on the audio interruption parameter, the instruction issuance standard for issuing the interception analysis instruction for the audio information according to the mouth aspect ratio of the user's mouth features is adjusted to a corresponding value;

[0124] If the audio interruption parameter is less than or equal to the preset interruption parameter, the bit rate of the audio information transmitted by the information acquisition module is adjusted to a corresponding value based on the average value of the volume integral parameters of each fluctuation instruction.

[0125] Specifically, the preset interruption parameter D0 is selected within the interval [2, 4].

[0126] Specifically, when the number of fluctuation instructions is the maximum number within a single cycle, the processing parameters of each information obtained by the information collection module are adjusted based on the audio interruption parameter. The audio interruption parameter characterizes the interruption of the audio information. When the audio interruption parameter is greater than the preset interruption parameter, the audio information has a high-frequency interruption. At this time, due to the abnormality of the instruction issuance standard for intercepting and analyzing instructions, multiple audio information is still intercepted when the user is speaking. At this time, the instruction issuance standard for the analysis instruction is adjusted to improve the accuracy of effective audio information interception; when the audio interruption parameter is less than or equal to the preset interruption parameter, the information has a low-frequency interruption. At this time, the information is lost due to transmission abnormalities. At this time, by adjusting the bit rate parameter, the currently used bit rate value is increased, the audio data transmission speed is increased, and the sound interruption caused by data loss is reduced. While fully acquiring the text information of the meeting, the user's speech content is further accurately extracted.

[0127] Specifically, based on the audio interruption parameter, the instruction issuing standard for issuing the interception analysis instruction for the audio information according to the mouth aspect ratio of the user's mouth feature is adjusted to a corresponding value, wherein,

[0128] The increase in the command issuance standard is proportional to the audio interruption parameter.

[0129] In this embodiment, optionally,

[0130] Comparing the audio interruption parameter with the first preset interruption comparison parameter and the second preset interruption comparison parameter;

[0131] If the audio interruption parameter is less than or equal to the first preset interruption comparison parameter, the preset deformation parameter is adjusted to 1.11 times the initial preset deformation parameter;

[0132] If the audio interruption parameter is less than or equal to the second preset interruption comparison parameter and greater than the first preset interruption comparison parameter, the preset deformation parameter is adjusted to 1.21 times the initial preset deformation parameter;

[0133] If the audio interruption parameter is greater than the second preset interruption comparison parameter, the preset deformation parameter is adjusted to 1.31 times the initial preset deformation parameter;

[0134] The first preset interrupt comparison parameter is 1.5D0, and the second preset interrupt comparison parameter is 2D0.

[0135] Specifically, the bit rate of the audio information transmitted by the information acquisition module is adjusted to a corresponding value based on the average value of the volume integral parameter of each fluctuation instruction, wherein:

[0136] The increase in the bit rate is inversely proportional to the average value of the volume integral parameter of each fluctuation instruction.

[0137] In this embodiment, optionally,

[0138] Recording the average value of the volume integral parameters of each fluctuation instruction as the average integral parameter, and comparing the average integral parameter with the first preset integral comparison threshold and the second preset integral comparison threshold;

[0139] If the average integral parameter is less than or equal to the first preset integral comparison threshold, the bit rate of the audio information transmitted by the information acquisition module is adjusted to 1.28 times the initial bit rate;

[0140] If the average integral parameter is less than or equal to the second preset integral comparison threshold and greater than the first preset integral comparison threshold, the bit rate of the audio information transmitted by the information acquisition module is adjusted to 1.19 times the initial bit rate;

[0141] If the average integral parameter is greater than the second preset integral comparison threshold, the bit rate of the audio information transmitted by the information acquisition module is adjusted to 1.12 times the initial bit rate;

[0142] The first preset integral comparison threshold is 0.64Y0, and the second preset integral comparison threshold is 0.8Y0.

[0143] Specifically, based on the average value of the verification parameters of each delay instruction, the time node corresponding to the start analysis instruction is preset in advance for the acquisition time, wherein,

[0144] The increase of the determined preset acquisition time length is proportional to the average value of the verification parameters of each delay instruction.

[0145] In this embodiment, optionally,

[0146] Recording the average value of the verification parameters of each delay instruction as the average verification parameter, and comparing the average verification parameter with a first preset verification comparison threshold and a second preset verification comparison threshold;

[0147] If the average verification parameter is less than or equal to the first preset verification comparison threshold, the preset acquisition time is adjusted to 1.13 times the initial preset acquisition time;

[0148] If the average verification parameter is less than or equal to the second preset verification comparison threshold and greater than the first preset verification comparison threshold, the preset acquisition time is adjusted to 1.23 times the initial preset acquisition time;

[0149] If the average verification parameter is greater than the second preset verification comparison threshold, the preset acquisition time is adjusted to 1.28 times the initial preset acquisition time;

[0150] The first preset verification comparison threshold is 1.6A2, and the second preset verification comparison threshold is 2.3A2.

[0151] Specifically, when the number of delay instructions is the maximum number within a single cycle, in this case, the intercepted audio information is incomplete due to the presence of a large number of interception and analysis instruction delays. In this case, the acquisition time of the time node corresponding to the start analysis instruction is preset in advance, and the preset acquisition time is further determined based on the verification parameters of each delay instruction. While accurately determining the specific circumstances of each information acquisition, the efficiency of extracting the conference content is further improved.

[0152] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

[0153] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A conference assistance system based on information analysis, characterized in that: include: An information collection module, which is used to store the video information and audio information of each user synchronously acquired; A feature extraction module connected to the information acquisition module, for extracting the mouth features of each user based on the acquired video information, and issuing an interception and analysis instruction for the audio information according to the mouth aspect ratio of the user's mouth features; the interception and analysis instruction includes a start analysis instruction and a stop analysis instruction, and the feature extraction module issues a start analysis instruction for the audio information under the condition that the mouth aspect ratio of the user's mouth features is greater than a preset deformation parameter; Obtaining the duration when the aspect ratio of the mouth is less than or equal to the preset deformation parameter, and issuing a termination analysis instruction for the audio information when the duration is greater than the preset holding time; Determine the duration between the time point when the mouth aspect ratio of the user's mouth feature is greater than the preset deformation parameter and the time point when the termination analysis instruction is issued for the audio information as the feature change duration; An audio extraction module, which is connected to the feature extraction module and the information collection module, and is used to intercept audio information based on the interception analysis instruction, and convert the intercepted audio information into text to obtain the conference text content; for a single interception analysis instruction, the audio extraction module intercepts the audio information within the corresponding time period based on the interception analysis instruction, and determines the duration of the audio within the corresponding time period; A calibration division module, which is connected to the audio extraction module and the information acquisition module, is used to determine a verification parameter and a volume integral parameter based on the audio information for a single interception and analysis instruction, and to perform type calibration on each interception and analysis instruction based on the verification parameter and the volume integral parameter; a volume time domain curve is drawn based on the volume of the audio information corresponding to the single interception and analysis instruction, and the integral value of the calculated volume time domain curve is recorded as the volume integral parameter; the calibration division module is used to calculate the ratio of the feature change duration to the audio duration, and record it as the verification parameter; An analysis module is connected to the calibration division module, the information acquisition module, the feature extraction module and the audio extraction module, and is used to periodically determine whether the acquisition of information is qualified based on the number of each type of interception and analysis instructions that are statistically calibrated. When it is determined that the acquisition of information is abnormal, the instruction issuance standard for issuing the interception and analysis instruction is adjusted to a corresponding value, the bit rate of the audio information transmitted by the information acquisition module is adjusted to a corresponding value, or the time node corresponding to the starting analysis instruction of the interception and analysis instruction is preset in advance for the acquisition duration.

2. The conference assistance system based on information analysis according to claim 1, characterized in that: The calibration and classification module is used to calibrate the interception and analysis instructions based on the verification parameters, including: If the verification parameter is less than or equal to the first preset verification parameter, the intercepted analysis instruction is marked as a qualified instruction; If the verification parameter is less than or equal to the second preset verification parameter and greater than the first preset verification parameter, the type of the interception analysis instruction is calibrated based on the interception duration of the corresponding time period intercepted by the interception analysis instruction; If the verification parameter is greater than the second preset verification parameter, the type of the interception analysis instruction is calibrated based on the volume integral parameter.

3. The conference assistance system based on information analysis according to claim 2, characterized in that: The calibration and classification module is used to calibrate the type of interception and analysis instructions based on the interception duration, including: If the interception time is less than or equal to the preset interception time, the type of the interception analysis instruction is calibrated based on the volume integral parameter; If the interception duration is longer than the preset interception duration, the interception analysis instruction will be marked as a low-sensitivity instruction.

4. The conference assistance system based on information analysis according to claim 3 is characterized in that: The calibration division module is used to calibrate the type of interception analysis instruction based on the volume integral parameter, including: If the volume integral parameter is less than or equal to the preset volume integral parameter, the interception analysis instruction is marked as a fluctuation instruction; If the volume integral parameter is greater than the preset volume integral parameter, the intercepted analysis instruction is marked as a delay instruction.

5. The conference assistance system based on information analysis according to claim 4 is characterized in that: The analysis module is used to determine whether the acquisition of information is qualified based on the statistical quantity of each type of interception analysis instructions marked, including: If the number of qualified instructions is the maximum number in a single cycle, it is determined that the acquisition of information is qualified, and the video information and audio information are continuously analyzed; If the number of low-sensitivity instructions is the maximum number in a single cycle, it is determined that the acquisition of information is abnormal, and the first preset check parameter and the second preset check parameter are adjusted based on the variance of the interception time length of each low-sensitivity instruction; If the number of fluctuation instructions is the maximum number in a single cycle, it is determined that the acquisition of information is abnormal, the number of intersections between the volume time domain curve of a single fluctuation instruction and the time axis is obtained, and the average number of intersections of each fluctuation instruction calculated is recorded as the audio interruption parameter, and the processing parameters of each information obtained by the information acquisition module are adjusted based on the audio interruption parameter; If the number of delay instructions is the maximum number in a single cycle, it is determined that the acquisition of information is abnormal, and the time node corresponding to the start analysis instruction is preset in advance for the acquisition duration based on the average value of the verification parameters of each delay instruction.

6. The conference assistance system based on information analysis according to claim 5, characterized in that: The analysis module is used to adjust the first preset check parameter and the second preset check parameter based on the variance of each interception time length of each low-sensitivity instruction, including: If the variance is less than or equal to the preset variance, adjusting the first preset check parameter and the second preset check parameter to corresponding values ​​based on the average value of each interception duration of each low-sensitivity instruction; If the variance is greater than the preset variance, the first preset verification parameter and the second preset verification parameter of the user corresponding to each low-sensitivity instruction are adjusted to corresponding values ​​based on the average value of each interception time length of each low-sensitivity instruction.

7. The conference assistance system based on information analysis according to claim 6, characterized in that: The analysis module is used to adjust the first preset check parameter and the second preset check parameter to corresponding values ​​based on the average value of each interception time length of each low-sensitivity instruction; The reduction range of the first preset check parameter and the second preset check parameter is proportional to the average value of each interception time length of each low-sensitivity instruction.

8. The conference assistance system based on information analysis according to claim 7, characterized in that: The analysis module is used to adjust the processing parameters of each information acquired by the information acquisition module based on the audio interruption parameter, including: If the audio interruption parameter is greater than the preset interruption parameter, then based on the audio interruption parameter, the instruction issuance standard for issuing the interception analysis instruction for the audio information according to the mouth aspect ratio of the user's mouth features is adjusted to a corresponding value; If the audio interruption parameter is less than or equal to the preset interruption parameter, the bit rate of the audio information transmitted by the information acquisition module is adjusted to a corresponding value based on the average value of the volume integral parameters of each fluctuation instruction.

9. The conference assistance system based on information analysis according to claim 8, characterized in that: The analysis module is used to adjust the instruction issuance standard for issuing the interception analysis instruction for the audio information according to the mouth aspect ratio of the user's mouth feature to a corresponding value based on the audio interruption parameter, wherein, The increase in the command issuance standard is proportional to the audio interruption parameter.

10. The conference assistance system based on information analysis according to claim 9, characterized in that: The analysis module is used to adjust the bit rate of the audio information transmitted by the information acquisition module to a corresponding value based on the average value of the volume integral parameter of each fluctuation instruction, wherein: The increase in the bit rate is inversely proportional to the average value of the volume integral parameter of each fluctuation instruction; Based on the average value of the verification parameters of each delay instruction, the time node corresponding to the start analysis instruction is preset in advance for the acquisition time, wherein, The increase of the determined preset acquisition time length is proportional to the average value of the verification parameters of each delay instruction.

Citation Information

Patent Citations

  • Conference content analysis method and system based on artificial intelligence and conference all-in-one machine

    CN118660128A

  • Method for synchronously calibrating audio and video streams and related product thereof

    CN115250373A

  • Audio and video positioning system for conference room scene

    CN118301279A