Unstructured business data interaction method and system based on large language model

By using audio feature analysis and filtering, the problem of poor interactive effects caused by noise interference in voice data in the aviation field has been solved, achieving efficient and accurate unstructured data interaction and improving the accuracy of text information and the quality of feedback.

CN122392537APending Publication Date: 2026-07-14BEIJING DEXUN AVIATION SERVICE CO LTD
View PDF 0 Cites -1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING DEXUN AVIATION SERVICE CO LTD
Filing Date
2026-05-29
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

In the aviation field, background noise interference in voice data leads to poor interaction of unstructured data based on large language models, resulting in large errors in text information recognition and understanding.

Method used

By acquiring the audio signal to be processed and historical reference audio signals, audio feature analysis is performed to determine the speech retention rate. Based on the speech retention rate, filtering is performed to remove noise interference, retain effective speech information, and then the information is converted into text information and input into a large language model for feedback.

Benefits of technology

It significantly improves the accuracy of text information, ensuring that the large language model can accurately parse user needs, output feedback results that meet expectations, and improve the interaction effect of unstructured business data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122392537A_ABST
    Figure CN122392537A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, in particular to an unstructured business data interaction method and system based on a large language model, which solves the technical problem of poor interaction effect caused by recognition errors in the prior art. The method comprises: obtaining a to-be-processed audio signal carrying user speech and at least one historical reference audio signal; performing audio feature analysis on the audio signals of each time period in the to-be-processed audio signal based on the at least one historical reference audio signal, and determining the speech retention degree corresponding to the audio signals of each time period in the to-be-processed audio signal; the speech retention degree is used to represent the attention weight of the audio signals of the corresponding time period in the filtering process; filtering the to-be-processed audio signal according to the speech retention degree corresponding to the audio signals of each time period to obtain a target speech signal; converting the target speech signal into text information, inputting the text information into a large language model, and outputting a feedback result for the text information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a method and system for unstructured business data interaction based on a large language model. Background Technology

[0002] In the aviation industry, unstructured business data interaction based on large language models leverages the natural language understanding, multimodal parsing, and reasoning capabilities of large language models to enable the intelligent flow, parsing, and interaction of unstructured data such as document texts, voice dialogues, images, and videos between airlines, airports, regulatory agencies, and various business systems in aviation scenarios. This solves the problems of low efficiency and large errors in traditional manual processing, covering multiple core scenarios such as cargo, customer service, and operations and maintenance.

[0003] In the service collaboration between airlines and passengers, the audio in voice or video messages sent by passengers is easily interfered with by various background noises. These noises may come from aviation-related environments such as the airport and cabin, or from everyday environmental sounds. As a result, the recognition and understanding of textual information in the voice message can be quite erroneous, leading to poor interactive performance based on large language models. Summary of the Invention

[0004] To address the technical problem of poor interaction performance caused by recognition errors in existing technologies, the present invention aims to provide a method and system for unstructured business data interaction based on a large language model. The specific technical solution adopted is as follows: This application provides a method for unstructured business data interaction based on a large language model, including: Acquire the audio signal to be processed carrying the user's voice and at least one historical reference audio signal; Based on at least one historical reference audio signal, audio feature analysis is performed on the audio signals of each time period in the audio signal to be processed to determine the speech preservation degree corresponding to the audio signal of each time period in the audio signal to be processed; the speech preservation degree is used to characterize the attention weight of the audio signal of the corresponding time period in the filtering process. The audio signal to be processed is filtered according to the speech preservation level of the audio signal in each time period to obtain the target speech signal; The target speech signal is converted into text information, and the text information is input into a large language model, which then outputs feedback results for the text information.

[0005] This application provides an unstructured business data interaction system based on a large language model, including: The signal acquisition unit is used to acquire the audio signal to be processed carrying the user's voice and at least one historical reference audio signal; The feature analysis unit is used to perform audio feature analysis on the audio signals of each time period in the audio signal to be processed based on at least one historical reference audio signal, and to determine the speech preservation degree corresponding to the audio signal of each time period in the audio signal to be processed; the speech preservation degree is used to characterize the attention weight of the audio signal of the corresponding time period in the filtering process. The filtering unit is used to filter the audio signal to be processed according to the speech preservation degree corresponding to the audio signal in each time period to obtain the target speech signal; The interactive execution unit is used to convert the target speech signal into text information, input the text information into the large language model, and output feedback results for the text information.

[0006] The present invention has the following beneficial effects: To address the issue of poor interactive performance in existing technologies due to recognition errors, this application provides a method and system for interacting with unstructured business data based on a large language model. This method acquires the audio signal to be processed, carrying user speech, and at least one historical reference audio signal. Using the historical reference audio signal as an analysis benchmark, audio feature analysis is performed on the audio signals of each time period in the audio signal to be processed. This determines the speech retention rate corresponding to each time period, thereby accurately identifying the effective and ineffective speech periods in the audio signal. Based on the speech retention rate of each time period, the audio signal is then filtered to obtain the target speech signal. This makes the filtering process more targeted, effectively removing noise interference while maximizing the retention of effective speech information. Finally, this application converts the target speech signal into text information, inputs the text information into a large language model, and outputs feedback results for the text information. This approach significantly improves the accuracy of the text information, ensuring that the large language model can accurately interpret user needs and output expected feedback results, thereby improving the interactive effect of unstructured business data and meeting the needs for efficient and accurate interaction in business scenarios. Attached Figure Description

[0007] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0008] Figure 1 A system architecture diagram of an unstructured business data interaction system based on a large language model, provided as an embodiment of the present invention; Figure 2This is a flowchart illustrating a method for unstructured business data interaction based on a large language model, provided as an embodiment of the present invention. Detailed Implementation

[0009] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of the unstructured business data interaction method and system based on a large language model proposed in this invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0010] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0011] The specific solution of the unstructured business data interaction method and system based on a large language model provided by the present invention will be described in detail below with reference to the accompanying drawings.

[0012] Please see Figure 1 The diagram illustrates a system architecture of an unstructured business data interaction system based on a large language model according to an embodiment of the present invention. The unstructured business data interaction system 10 based on a large language model includes: a signal acquisition unit 11, a feature analysis unit 12, a filtering processing unit 13, and an interaction execution unit 14.

[0013] The signal acquisition unit 11 is used to acquire the audio signal to be processed carrying the user's voice and at least one historical reference audio signal.

[0014] For example, the signal acquisition unit 11 can employ an audio acquisition device (such as a high-fidelity microphone or audio / video recorder) or a data interface module (for receiving audio data uploaded by the user through the terminal). The acquired or received audio signals to be processed include pure audio signals sent by the user during business interactions, audio streams in audio / video signals, etc. Acquisition parameters can be set according to the business scenario to ensure the basic quality of the audio signal. Historical reference audio signals are audio samples representing historical business interactions. For example, historical reference audio signals originate from historical business interaction audio stored in the business database.

[0015] The feature analysis unit 12 is used to perform audio feature analysis on the audio signals of each time period in the audio signal to be processed based on at least one historical reference audio signal, and to determine the speech preservation degree corresponding to the audio signal of each time period in the audio signal to be processed.

[0016] Among them, speech preservation is used to characterize the attention weight of the audio signal in the corresponding time period during the filtering process.

[0017] For example, the feature analysis unit 12 can integrate a data processing chip or a software module deployed in a server. It can distinguish the time period corresponding to effective speech and the time period corresponding to background noise by comparing the feature correlation between the audio signal to be processed and the historical reference audio signal, and quantify the importance of each time period. Finally, it outputs the speech preservation parameter, providing a precise basis for subsequent filtering processing.

[0018] The filtering processing unit 13 is used to filter the audio signal to be processed according to the speech retention degree corresponding to the audio signal in each time period to obtain the target speech signal.

[0019] For example, the filtering processing unit 13 can integrate a speech filtering model (such as a Long Short-Term Memory (LSTM) network model). By constructing a weighted mask matrix that matches the feature dimensions of the audio signal to be processed, it performs differentiated filtering on time periods with different retention levels. It focuses on retaining effective speech information in time periods with high retention levels and strengthens noise filtering in time periods with low retention levels, thereby removing noise while retaining key speech content in business interactions to the greatest extent.

[0020] The interactive execution unit 14 is used to convert the target speech signal into text information, input the text information into the large language model, and output feedback results for the text information.

[0021] For example, the interactive execution unit 14 integrates an Automatic Speech Recognition (ASR) module, an Optical Character Recognition (OCR) module, and a large language model interface. The ASR module employs a deep learning-based speech recognition algorithm (such as a Transformer architecture recognition model), the OCR module uses an integrated text detection and recognition model, and the large language model interface can connect to large vertical domain models adapted to airline customer service business scenarios. It can parse user needs based on text information, generate feedback results (such as business processing receipts and consultation replies) and execution instructions (such as calling business system interfaces and transferring to human customer service).

[0022] It should be noted that the various embodiments of this application can be referenced or learned from each other. For example, the same or similar steps, method embodiments, system embodiments and device embodiments can be referenced from each other without limitation.

[0023] Please see Figure 2The diagram illustrates a flowchart of an embodiment of the present invention for a method of unstructured business data interaction based on a large language model, the method comprising the following steps: Step 201: Obtain the audio signal to be processed carrying the user's voice and at least one historical reference audio signal.

[0024] The audio signals to be processed are audio data generated by users during business interactions that require interactive processing, including but not limited to pure audio uploaded by users through voice customer service calls and terminals, as well as audio streams in audio and video files. Historical reference audio signals are audio samples representing historical business interactions, including audio samples that occurred historically in the business scenario and are related to the business interaction, such as valid voice segments from historical customer service consultations and key voice records during business processing. Their purpose is to provide a reference benchmark for the feature analysis of the audio signals to be processed.

[0025] In some embodiments, after acquiring the audio signal to be processed, it can be preprocessed, including format standardization (e.g., uniformly converting to WAV format) and sampling rate standardization (e.g., adjusting to 16kHz) to ensure the consistency of subsequent analysis. Historical reference audio signals need to be screened from the business database, and the screening criteria are related to the current business scenario (e.g., in the airline scenario, historical audio related to flight consultation and rescheduling is screened), and after preliminary denoising and invalid segment removal, irrelevant data is avoided from interfering with the analysis results.

[0026] Step 202: Based on at least one historical reference audio signal, perform audio feature analysis on the audio signals of each time period in the audio signal to be processed, and determine the speech preservation degree corresponding to the audio signal of each time period in the audio signal to be processed.

[0027] Among them, speech preservation is used to characterize the attention weight of the audio signal in the corresponding time period during the filtering process.

[0028] This application can distinguish between time periods containing valid business speech and time periods dominated by noise and invalid speech in the audio signal to be processed through audio feature analysis, and determine the importance of each time period through quantitative indicators, namely speech retention. The speech retention value can be set from 0 to 1. The closer the speech retention value is to 1, the more important the audio signal of that time period is, and the more important it should be to be preserved during filtering. The closer the speech retention value is to 0, the more important the audio signal of that time period is, and the more important it should be to be preserved during filtering. The closer the speech retention value is to 0, the more important the audio signal of that time period is, and the more important it should be to be suppressed during filtering.

[0029] Step 203: Filter the audio signal to be processed according to the speech retention rate of the audio signal in each time period to obtain the target speech signal.

[0030] In one possible implementation, this application can extract the speech feature sequence of the audio signal to be processed, and construct a weighted mask matrix that matches the dimension of the speech feature sequence based on the speech preservation degree of the audio signal at each time period of the audio signal to be processed.

[0031] The speech feature sequence can be represented by a Mel-Frequency Cepstral Coefficients (MFCC) feature matrix or a Mel Filter Bank Features (Fbank) feature matrix, which can effectively characterize the acoustic properties of the speech signal. The weighted mask matrix has the same dimension as the speech feature sequence, and the element values ​​in the weighted mask matrix are the speech preservation degree of the corresponding time period.

[0032] For example, a speech feature sequence can be represented as a two-dimensional matrix, where its time dimension corresponds to the time segment division of the audio signal. This application can map the speech retention of each time segment to all time frames within that time segment, thereby forming a weighted sequence with the same length as the time dimension of the speech feature sequence. Then, through a copying operation, this weighted sequence is aligned dimensionally with the two-dimensional structure of the speech feature sequence, forming the final weighted mask matrix. Thus, the dimension of this weighted mask matrix matches the speech feature sequence, and its element values ​​reflect the speech retention of the corresponding time segment.

[0033] Then, the speech feature sequence is weighted and the weighted mask matrix is ​​weighted to obtain the weighted speech feature sequence.

[0034] The weighting operation refers to multiplying the corresponding elements of the speech feature sequence with the corresponding elements of the weighted mask matrix element by element, using the elements in the weighted mask matrix as weighting coefficients, to obtain the weighted speech feature sequence. Thus, for periods with high retention, the feature values ​​remain at a high level, preserving effective speech features; for periods with low retention, the feature values ​​are suppressed, weakening noise features.

[0035] In this way, the weighted speech feature sequence can be input into the speech filtering model to obtain the target speech signal.

[0036] The speech filtering model can employ an LSTM model, which effectively captures the temporal dependencies of speech signals. During the training phase, the input to the LSTM model is a weighted speech feature sequence (obtained by weighting a noisy speech feature sequence with a corresponding weighted mask matrix), and the supervision signal is the corresponding clean speech signal. The training objective is to minimize the error between the model output and the clean speech signal. In the application phase (i.e., during filtering), the weighted speech feature sequence is input into the trained speech filtering model to obtain the target speech signal.

[0037] Step 204: Convert the target speech signal into text information, input the text information into the large language model, and output the feedback result for the text information.

[0038] The feedback results include business responses and execution instructions (such as calling the business system to handle business).

[0039] In one possible implementation, when the target speech signal originates from pure audio data, this application can convert the target speech signal into text information using an automatic speech recognition algorithm.

[0040] When the target speech signal originates from audio and video data, this application can convert the target speech signal into first text information using an automatic speech recognition algorithm, and perform optical character recognition on the image frames in the audio and video data to obtain second text information, and then merge the first text information and the second text information as the text information.

[0041] The pure audio data includes recordings of user calls to voice customer service and independent audio files uploaded by the terminal. The audio and video data includes video files uploaded by users that contain both voice and images (such as videos recorded during business transactions or videos of customer feedback). This application can process image frames in the audio and video data, including image preprocessing (such as grayscale conversion and noise reduction), text detection, and text recognition, to obtain second text information (such as document information and itinerary information in the image). When merging the first and second text information, this application can remove duplicate content in the text (such as a flight number mentioned in the audio matching the flight number displayed in the image, retaining only one), and supplement missing information (such as dates not clearly stated in the audio, extracted from the itinerary in the image), ensuring that the merged text information is comprehensive and accurate.

[0042] For example, in the audio and video data uploaded by the user, the audio content is "I want to change this flight", and the image frame displays the itinerary information "Flight number MU5101, date 2024-10-05". Then the first text information is "I want to change this flight", the second text information is "Flight number MU5101, date 2024-10-05", and the combined text information is "I want to change the flight number MU5101, date 2024-10-05", which provides a complete basis for the large language model to accurately parse the requirements.

[0043] Based on the above technical solution, this application acquires the audio signal to be processed carrying user speech and at least one historical reference audio signal. Using the historical reference audio signal as the analysis benchmark, it performs audio feature analysis on the audio signals of each time period in the audio signal to be processed, determining the speech retention rate corresponding to the audio signal of each time period in the audio signal to be processed. This accurately identifies the effective and ineffective speech periods in the audio signal to be processed. Then, based on the speech retention rate corresponding to the audio signal of each time period, the audio signal to be processed is filtered to obtain the target speech signal. This makes the filtering process more targeted, effectively removing noise interference while preserving effective speech information to the greatest extent. Finally, this application can convert the target speech signal into text information and input the text information into a large language model, outputting feedback results for the text information. The above solution significantly improves the accuracy of the text information, ensuring that the large language model can accurately parse user needs and output expected feedback results, thereby improving the interaction effect of unstructured business data and meeting the needs of efficient and accurate interaction in business scenarios.

[0044] As a possible embodiment of this application, step 202 above can be implemented through the following steps: Step 301: Perform time-series feature matching between each unit time period of the audio signal to be processed and each unit time period of each historical reference audio signal to obtain the time-series feature matching result.

[0045] The temporal feature matching result is used to characterize the correlation between each unit time segment of the audio signal to be processed and each unit time segment of at least one historical reference audio signal. A unit time segment refers to the smallest analytical unit into which the audio signal is divided according to a fixed duration; for example, a unit time segment can be set to 10 milliseconds. This duration ensures the precision of the temporal features while avoiding excessive computation. Temporal features can be characterized by short-time energy spectrum, which is a fundamental acoustic feature of speech signals. Changes in temporal features are highly correlated with prosody, pauses, and stress at the phonological level, reflecting the temporal variation patterns of speech and being crucial for determining the correlation between two audio segments.

[0046] For example, temporal feature matching can employ the Dynamic Time Warping (DTW) algorithm. By scaling and shifting the audio signal along the time axis, the DTW distance between a unit time interval of the audio signal to be processed and that of a unit time interval of the historical reference audio signal is calculated. A smaller DTW distance indicates higher feature similarity and a stronger correlation. The temporal feature matching result can be represented as a DTW distance matrix (or a tensor). The element values ​​in this DTW distance matrix represent the DTW distance between the corresponding unit time interval in the audio signal to be processed and the corresponding unit time interval in the historical reference audio signal.

[0047] Step 302: Based on the temporal feature matching results, divide the audio signal to be processed into audio signals of multiple time periods.

[0048] The process involves dividing the audio signal into voice segments and background segments. Voice segments refer to those segments in the audio signal to be processed that have a high correlation with historical reference audio signals and are likely to contain valid business voice. Background segments refer to those segments that have a low correlation with historical reference audio signals and are primarily composed of noise or invalid voice. This application can further divide the segments by setting a difference threshold.

[0049] For example, the difference threshold can be set to 0.8 times the average DTW distance, which can be calculated by averaging the DTW distances between each unit time segment of the audio signal to be processed and all unit time segments of all historical reference audio signals. For a unit time segment in the audio signal to be processed, if the DTW distance between one (or more) unit time segments of the historical reference audio signals and that unit time segment is less than the difference threshold, then that unit time segment is classified into the speech segment. Conversely, if the DTW distances between all unit time segments of the historical reference audio signals and that unit time segment are greater than or equal to the difference threshold, then that unit time segment is classified into the background segment.

[0050] Step 303: Perform audio feature analysis on the audio signal of each time period in the audio signal to be processed, and determine the speech preservation degree corresponding to the audio signal of each time period in the audio signal to be processed.

[0051] This application allows for separate processing operations for the speech segment and the background segment. For example, the speech preservation level of the audio signal in the background segment can be set to a preset minimum value, and further audio feature analysis can be performed only on the speech segment, thereby improving processing efficiency.

[0052] For example, this application can quantify the corresponding speech retention by analyzing audio feature dimensions such as business criticality (e.g., whether it contains core requirements) and effective information content (e.g., the degree of noise interference). Audio feature analysis can be combined with indicators such as signal energy and principal component features. For example, if the short-term energy mean of the speech segment is higher than that of the background segment, and the effective information is concentrated in a few principal component components, then the retention of the speech segment is higher.

[0053] Based on the above technical solution, this application establishes the correlation between the audio signal to be processed and the historical reference audio signal through temporal feature matching, thereby accurately distinguishing between speech periods and background periods, providing a clear analysis object for subsequent speech retention calculation. Through targeted feature analysis of different periods, the quantification of speech retention becomes more reasonable and accurate, further improving the effect of subsequent filtering processing, ensuring the quality of the target speech signal, and laying a good foundation for subsequent text conversion and large language model interaction.

[0054] As a possible embodiment of this application, step 302 above can be implemented through the following steps: Step 401: Based on the timing feature matching results, determine the cumulative matching value for each unit time period in the audio signal to be processed.

[0055] Among them, the cumulative matching value is used to characterize the number of times a unit time period is determined to be a match in the time series feature matching results.

[0056] For example, this application can determine the DTW distance between each unit time segment in the audio signal to be processed and each unit time segment of each historical reference audio signal based on the time sequence feature matching result, and take the number of DTW distances less than the matching threshold (which can be adjusted according to the actual situation, for example, it can be set to 0.8 times the average of all DTW distances with each historical reference audio signal) as the cumulative matching value of the unit time segment.

[0057] In one possible implementation, this application can filter historically associated audio signals from at least one historical reference audio signal based on the temporal feature matching results, and determine the cumulative matching value of each unit time period in the audio signal to be processed based on the temporal feature matching results of each unit time period in the audio signal to be processed and each unit time period in the historically associated audio signal.

[0058] For example, this application can initially filter historically associated audio signals related to the audio signal to be processed from at least one historical reference audio signal based on the temporal feature matching results. For instance, for a historical reference audio signal, the average DTW distance between each unit time period in the historical reference audio signal and each unit time period in the audio signal to be processed can be defined as the overall DTW distance of the historical reference audio signal, used to measure the overall similarity between the historical reference audio signal and the audio signal to be processed. Subsequently, historically associated audio signals are obtained by filtering the overall DTW distance of each historical reference audio signal (e.g., the overall DTW distance is less than 0.7 times the average of the overall DTW distances of all historical reference audio signals). In this way, this application can calculate the cumulative matching value based solely on historically associated audio signals, avoiding matching interference and computational redundancy caused by historical reference audio signals unrelated to the audio signal to be processed.

[0059] If the number of historical associated audio signals after screening is too small (e.g., less than 2), the association threshold can be appropriately lowered, or historical reference audio signals that are highly related to the business scenario of the audio signal to be processed can be added to ensure the reliability of the cumulative value calculation. If the number of historical associated audio signals is too large (e.g., more than 10), they can be further sorted according to the overall similarity, and the top few (e.g., the top 8) historical associated audio signals with the highest similarity can be selected to participate in the calculation to balance the calculation efficiency and accuracy.

[0060] Step 402: Based on the cumulative matching value, divide the continuous unit time period with a cumulative matching value of zero into background time period, and divide the continuous unit time period with a cumulative matching value of non-zero into speech time period.

[0061] A cumulative match value of zero indicates that the unit time period has no effective correlation with any historical reference audio signals, and is likely background noise or invalid speech. Consecutive units of this type constitute background time periods. A cumulative match value of non-zero indicates that the unit time period has a valid correlation with at least one historical reference audio signal, and is highly likely to contain valid business speech. Consecutive units of this type constitute speech time periods.

[0062] In some embodiments, to avoid fragmentation of time period division due to matching anomalies in individual unit time periods, a minimum length threshold for consecutive unit time periods can be set. For example, the minimum length of the background time period can be set to 5 unit time periods (50 milliseconds), and the minimum length of the voice time period can be set to 10 unit time periods (100 milliseconds). If consecutive unit time periods are shorter than the minimum length, they can be merged into adjacent main time periods to ensure the integrity and practicality of time period division.

[0063] Based on the above technical solution, this application determines the cumulative matching value of each unit time period in the audio signal to be processed according to the time sequence feature matching result, thereby quantifying the correlation strength between each unit time period and the historical reference audio signal. Then, according to the cumulative matching value, the continuous unit time periods with a cumulative matching value of zero are divided into background time periods, and the continuous unit time periods with a cumulative matching value of non-zero are divided into speech time periods, thereby accurately distinguishing between background time periods and speech time periods, reducing the interference of noise and invalid information on subsequent processing, and ensuring the integrity of effective speech time periods, providing a more reliable basis for subsequent speech preservation calculation and filtering processing.

[0064] As a possible embodiment of this application, step 303 above can be implemented through the following steps: Step 501: Based on the temporal feature matching results and the signal energy of the audio signal in each time period of the audio signal to be processed, determine the interactive key performance degree corresponding to each speech time period in the audio signal to be processed.

[0065] Among them, the interaction criticality is used to characterize the criticality of voice segments in business interactions.

[0066] In one possible implementation, this application can determine the speech criticality coefficient of each speech segment in the audio signal to be processed based on the temporal feature matching results.

[0067] Among them, the speech criticality coefficient is used to characterize the semantic significance of speech segments.

[0068] For example, this application can determine the importance coefficient and concentration coefficient corresponding to each speech segment in the audio signal to be processed based on the temporal feature matching result, and determine the speech criticality coefficient based on the importance coefficient and concentration coefficient corresponding to each speech segment in the audio signal to be processed.

[0069] The importance coefficient satisfies the following formula: in, This represents the importance coefficient corresponding to this speech segment. This is the average of the cumulative matching values ​​for all unit time periods within the given speech time period. This represents the maximum cumulative matching value across all time units within the given speech period. This represents the maximum value of the cumulative matched values ​​across all time segments in the audio signal to be processed. This represents the duration of the speech segment.

[0070] in, pass Will Normalized to between 0 and 1 to eliminate differences in magnitude, and The importance coefficient is obtained by multiplying the values. This importance coefficient is positively correlated with the cumulative matching value and the duration. That is, the higher the cumulative matching value and the longer the duration, the larger the importance coefficient, indicating that the core voice information of the voice segment is richer and the basic importance in business interaction is higher.

[0071] For the first speech segment in the audio signal to be processed, the centrality coefficient satisfies the following formula: in, This is the centrality coefficient corresponding to this speech period. This represents the importance coefficient corresponding to the next speech segment adjacent to this speech segment. The duration of the interval between the next speech segment adjacent to this speech segment. For normalization functions (e.g., maximum-minimum normalization, where the maximum value is the value of each speech segment). The maximum value and minimum value are for each speech segment. The minimum value in the equation is used to map the calculation result to the range of 0 to 1.

[0072] For the middle speech segment in the audio signal to be processed, the centrality coefficient satisfies the following formula: in, This is the centrality coefficient corresponding to this speech period. This is the importance coefficient corresponding to the preceding speech segment adjacent to this speech segment. The duration of the interval between the preceding speech segment and the current speech segment. This represents the importance coefficient corresponding to the next speech segment adjacent to this speech segment. The duration of the interval between the next speech segment adjacent to this speech segment. For normalization functions (e.g., maximum-minimum normalization, where the maximum value is the value of each speech segment). The maximum value and minimum value are for each speech segment. The minimum value in the equation is used to map the calculation result to the range of 0 to 1.

[0073] For the last speech segment in the audio signal to be processed, the centrality coefficient satisfies the following formula: in, This is the centrality coefficient corresponding to this speech period. This is the importance coefficient corresponding to the preceding speech segment adjacent to this speech segment. The duration of the interval between the preceding speech segment and the current speech segment. For normalization functions (e.g., maximum-minimum normalization, where the maximum value is the value of each speech segment). The maximum value and minimum value are for each speech segment. The minimum value in the equation is used to map the calculation result to the range of 0 to 1.

[0074] and The ratios of the core importance of adjacent speech segments to their interval lengths are respectively represented. The larger the ratio, the more prominent the importance of the adjacent speech segments is relative to the interval. The average of the two ratios, after normalization, yields a central tendency coefficient, which reflects the degree of clustering of the speech segment with adjacent important speech segments. The larger the value, the stronger the correlation between the voice segment and adjacent valid voice segments, the more concentrated the overall valid voice region, and the higher the semantic coherence and completeness in business interactions.

[0075] When only one speech segment is identified in the audio signal to be processed, since there are no adjacent speech segments, this application can directly set the concentration coefficient of the speech segment to a preset value, such as 1, indicating that the speech segment has independent concentration in time sequence.

[0076] For example, the speech criticality coefficient satisfies the following formula: in, This represents the speech criticality coefficient for that speech segment. This is the centrality coefficient corresponding to this speech period. This represents the importance coefficient corresponding to this speech segment. For normalization functions (e.g., maximum-minimum normalization, where the maximum value is the value of each speech segment). The maximum value and minimum value are for each speech segment. The minimum value in the equation is used to map the calculation result to the range of 0 to 1.

[0077] Subsequently, this application can determine the tone performance coefficient of each speech segment in the audio signal to be processed based on the signal energy of the audio signals of each speech segment in the audio signal to be processed and the background segments adjacent to the speech segments.

[0078] The tone performance coefficient is used to characterize the degree to which a user emphasizes their tone during a speech segment. Signal energy can be represented by the mean of short-time energy, which is the average energy of the speech signal within a unit of time and reflects the strength of the speech. When expressing important business needs, users often unconsciously emphasize their tone, resulting in a significantly higher mean of short-time energy during the speech segment compared to adjacent background segments.

[0079] For example, when there are adjacent background periods before and after the speech period, the tone performance coefficient satisfies the following formula: in, This represents the tone performance coefficient for that speech segment. This represents the average short-time energy across all time intervals within the given speech period. The background energy benchmark corresponding to this speech period satisfies , It is the average of the short-time energy of all unit time intervals within the preceding background time interval adjacent to the speech time interval. It is the average of the short-time energy of all unit time periods within the next background time period adjacent to the speech time period. This is a function to find the maximum value, used for... Negative values ​​are truncated to avoid negative differences due to excessive background noise.

[0080] In this formula, The absolute energy component reflects the energy intensity of the speech signal itself during a speech segment. The energy excess component is used to reflect the net energy contribution of the speech signal energy during a speech segment to the background energy reference. Determining the value by the arithmetic mean of the energy of adjacent background time periods can effectively eliminate the influence of local background energy fluctuations. Since the absolute energy component reflects the tone characteristics from the perspective of the intensity of the speech itself, and the energy excess component reflects the tone characteristics from the perspective of the prominence of the speech relative to the background noise, the two have a similar influence on tone evaluation. Therefore, the arithmetic mean can be used to fuse the two, and the fusion result can be mapped to the same order of magnitude range as the original components to ensure that the numerical range of the tone performance coefficient has physical consistency.

[0081] It should be noted that this application has already normalized the short-time energy during the short-time energy extraction process. This reflects the short-term energy difference between the speech segment and the local background environment. The greater the difference, the more obvious the emphasis on the tone. The tone performance coefficient of the speech segment is obtained by averaging the short-term energy of all unit segments within the speech segment. The larger the tone performance coefficient, the more obvious the user's emphasis on the tone within the speech segment, and the more attention the speech information of the speech segment receives in business interactions.

[0082] When the speech segment only has the adjacent preceding background segment, it is equivalent to The tone expression coefficient satisfies the following formula: in, This represents the tone performance coefficient for that speech segment. This is the average of the short-time energy across all unit time intervals within the speech duration. It is the average of the short-time energy of all unit time periods within the preceding background time period adjacent to the speech time period. This is a function to find the maximum value, used for... Negative values ​​are truncated to avoid negative differences due to excessive background noise.

[0083] When the speech segment only has an adjacent background segment, it is equivalent to The tone expression coefficient satisfies the following formula: in, This represents the tone performance coefficient for that speech segment. This is the average of the short-time energy across all unit time intervals within the speech duration. It is the average of the short-time energy of all unit time periods within the next background time period adjacent to the speech time period. This is a function to find the maximum value, used for... Negative values ​​are truncated to avoid negative differences due to excessive background noise.

[0084] When there is no background segment in the audio signal to be processed (i.e., the entire audio is determined to be a complete speech segment), it is equivalent to The tone expression coefficient satisfies the following formula: in, This represents the tone performance coefficient for that speech segment. This is the average of the short-time energy of all unit time intervals within the speech period. In other words, when there is no background time interval in the audio signal to be processed, this application can use the formula above... Set it to 0 to obtain the tone expression coefficient.

[0085] Compared to methods that simply use absolute energy thresholds or energy differences, this application employs a dual-component fusion mechanism that simultaneously considers both the absolute intensity and relative prominence of speech. For example, in high-noise environments such as airports, where background energy is high, using only energy differences might lead to underestimation of speech intensity when the user speaks loudly but the difference relative to the background is not prominent. Conversely, using only absolute energy might result in misjudgment when background noise occasionally intensifies. This application's dual-component fusion mechanism effectively mitigates the limitations of single indicators, improving the robustness and accuracy of tone emphasis assessment in complex aviation business interaction scenarios.

[0086] Thus, this application can determine the interactive key performance degree corresponding to each speech segment in the audio signal to be processed based on the speech keyness coefficient and tone performance coefficient of each speech segment in the audio signal to be processed.

[0087] For example, the key performance characteristics of an interaction satisfy the following formula: in, This refers to the key performance indicators of the interaction corresponding to this voice segment. This represents the tone performance coefficient for that speech segment. This represents the speech criticality coefficient for that speech segment. The speech criticality coefficient reflects the semantic criticality, while the tone performance coefficient reflects the expressive salience. The interaction criticality score is obtained by calculating the average value, which can comprehensively characterize the criticality of the speech segment in business interactions.

[0088] It should be noted that if there is no background speech segment in the audio signal to be processed, this application can use a clustering algorithm (such as the DBSCAN clustering algorithm) to divide the audio signal to be processed into multiple speech segments based on the cumulative matching value of each unit segment in the audio signal to be processed, and use the ratio of the mean of the cumulative matching value of all unit segments in each speech segment to the maximum value of the cumulative matching value of all unit segments in the audio signal to be processed as the interactive key performance of each speech segment.

[0089] Step 502: Perform principal component analysis on the audio signals of each speech segment in the audio signal to be processed to determine the effective component coefficients corresponding to each speech segment in the audio signal to be processed.

[0090] Among them, the effective component coefficient is used to characterize the effectiveness of the voice segment in business interaction. Principal component analysis decomposes the audio signal of the voice segment into multiple principal component components. By analyzing the eigenvalues ​​and feature similarities of each principal component component, the concentration and purity of effective voice information are quantified. The higher the effective component coefficient, the less noise interference the audio signal of that voice segment is affected and the higher the effective information content.

[0091] In one possible implementation, principal component analysis is performed on the audio signals of each speech segment in the audio signal to be processed to obtain multiple principal component components of the audio signals of each speech segment in the audio signal to be processed and the eigenvalues ​​corresponding to each principal component component.

[0092] Principal component analysis converts the audio signal of a speech segment (after framing and windowing) into multiple principal component components in the feature space. Each principal component component corresponds to an eigenvalue, and the magnitude of the eigenvalue reflects the effective information content contained in the principal component component. The larger the eigenvalue, the higher the effective information content.

[0093] Subsequently, feature distribution analysis was performed on the eigenvalues ​​corresponding to each principal component to determine the effective component coefficients for the corresponding speech time period.

[0094] For example, this application can first calculate the sum of eigenvalues ​​corresponding to all principal component components, and then set a cumulative contribution threshold based on the sum of eigenvalues ​​(e.g., 85% of the sum). Afterwards, each principal component is sorted in descending order of its corresponding eigenvalues, and the accumulated eigenvalues ​​are calculated one by one until the accumulated eigenvalue is first greater than or equal to the cumulative contribution threshold. Then, the number of accumulated eigenvalues ​​is determined. Furthermore, based on the number of such eigenvalues Determine the effective component coefficient.

[0095] The effective component coefficient satisfies the following formula: in, The effective component coefficients corresponding to this speech segment. The number of accumulated eigenvalues. This represents the number of eigenvalues ​​corresponding to each principal component. The smaller the value, the more concentrated the effective information is in the first few principal components. The larger the effective component coefficient, the less noise interference the audio signal of that speech segment is affected by, the more concentrated the effective information is, and the higher the signal quality.

[0096] Step 503: Based on the interactive key performance degree and effective component coefficient corresponding to each speech period in the audio signal to be processed, determine the speech retention degree corresponding to each speech period in the audio signal to be processed, and set the speech retention degree corresponding to each background period in the audio signal to be processed to a preset minimum value.

[0097] For example, the speech retention rate corresponding to the audio signal of a speech segment satisfies the following formula: in, This represents the speech retention rate of the audio signal corresponding to that speech period. The effective component coefficients corresponding to this speech segment. This represents the key performance level of the interaction corresponding to the speech segment. By using the mean of the effective component coefficient and the key performance level of the interaction as the speech retention rate, we ensure that the speech retention rate reflects both the keyness and effectiveness of the segment.

[0098] The preset minimum value can be set to 0, indicating that the audio signal in the background period is mainly noise and invalid information, which can be suppressed to the maximum extent during filtering.

[0099] Based on the above technical solutions, this application determines the key interactive performance of each speech segment according to the temporal feature matching results and the signal energy of the audio signal in each segment, thereby quantifying the business criticality of the speech segment. Principal component analysis is performed on the audio signals of each speech segment to determine the effective component coefficients corresponding to each speech segment, thereby quantifying the information effectiveness of the speech segment. Finally, the two are combined to calculate the speech retention rate, making the quantification of retention rate more comprehensive and accurate. In addition, the retention rate of the background segment is set to a preset minimum value, which clarifies the focus of the filtering process. This allows for the maximum retention of key and effective speech information while removing noise, further improving the quality of the target speech signal and providing a more reliable foundation for subsequent text conversion and large language model interaction.

[0100] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0101] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A method for unstructured business data interaction based on a large language model, characterized in that, include: Acquire the audio signal to be processed carrying the user's voice and at least one historical reference audio signal; Based on the at least one historical reference audio signal, audio feature analysis is performed on the audio signals of each time period in the audio signal to be processed to determine the speech retention degree corresponding to the audio signal of each time period in the audio signal to be processed; the speech retention degree is used to characterize the attention weight of the audio signal of the corresponding time period in the filtering process; The audio signal to be processed is filtered according to the speech retention rate corresponding to the audio signal in each time period to obtain the target speech signal; The target speech signal is converted into text information, and the text information is input into a large language model to output feedback results for the text information.

2. The method for unstructured business data interaction based on a large language model according to claim 1, characterized in that, Based on the at least one historical reference audio signal, audio feature analysis is performed on the audio signals of each time period in the audio signal to be processed to determine the speech preservation degree corresponding to the audio signal of each time period in the audio signal to be processed, including: The temporal feature matching is performed between each unit time period of the audio signal to be processed and each unit time period of each historical reference audio signal to obtain a temporal feature matching result; the temporal feature matching result is used to characterize the association relationship between each unit time period of the audio signal to be processed and each unit time period of the at least one historical reference audio signal. Based on the temporal feature matching results, the audio signal to be processed is divided into multiple time periods; the multiple time periods are divided according to speech time periods and background time periods. Audio feature analysis is performed on the audio signal of each time period in the audio signal to be processed to determine the speech preservation degree corresponding to the audio signal of each time period in the audio signal to be processed.

3. The method for unstructured business data interaction based on a large language model according to claim 2, characterized in that, Based on the temporal feature matching results, the audio signal to be processed is divided into multiple time-segment audio signals, including: Based on the temporal feature matching results, the cumulative matching value for each unit time period in the audio signal to be processed is determined; the cumulative matching value is used to characterize the number of times the unit time period is determined to have a match in the temporal feature matching results; Based on the cumulative matching value, consecutive time periods with a cumulative matching value of zero are divided into background time periods, and consecutive time periods with a cumulative matching value of non-zero are divided into speech time periods.

4. The method for unstructured business data interaction based on a large language model according to claim 3, characterized in that, Based on the temporal feature matching results, the cumulative matching value for each unit time period in the audio signal to be processed is determined, including: Based on the temporal feature matching results, historical associated audio signals are filtered from the at least one historical reference audio signal, and the cumulative matching value of each unit time period in the audio signal to be processed is determined based on the temporal feature matching results between each unit time period in the audio signal to be processed and each unit time period in the historical associated audio signal.

5. The method for unstructured business data interaction based on a large language model according to claim 2, characterized in that, Audio feature analysis is performed on the audio signal of each time segment in the audio signal to be processed to determine the speech preservation degree corresponding to the audio signal of each time segment in the audio signal to be processed, including: Based on the temporal feature matching results and the signal energy of the audio signal in each time period of the audio signal to be processed, the interaction key performance degree corresponding to each speech time period in the audio signal to be processed is determined; the interaction key performance degree is used to characterize the degree of importance of the speech time period in business interaction; Principal component analysis is performed on the audio signals of each speech segment in the audio signal to be processed to determine the effective component coefficients corresponding to each speech segment in the audio signal to be processed; the effective component coefficients are used to characterize the effectiveness of the speech segment in business interaction; Based on the interactive key performance degree and effective component coefficient corresponding to each speech segment in the audio signal to be processed, the speech retention degree corresponding to each speech segment in the audio signal to be processed is determined, and the speech retention degree corresponding to each background segment in the audio signal to be processed is set to a preset minimum value.

6. The method for unstructured business data interaction based on a large language model according to claim 5, characterized in that, Based on the temporal feature matching results and the signal energy of the audio signal in each time period of the audio signal to be processed, the interactive key performance degree corresponding to each speech time period in the audio signal to be processed is determined, including: Based on the temporal feature matching results, the speech criticality coefficient of each speech segment in the audio signal to be processed is determined; the speech criticality coefficient is used to characterize the semantic significance of the speech segment. Based on the signal energy of each speech segment in the audio signal to be processed and the audio signal of the background segment adjacent to the speech segment, the tone performance coefficient of each speech segment in the audio signal to be processed is determined; the tone performance coefficient is used to characterize the degree of emphasis of the user's tone within the speech segment. The interactive key performance degree corresponding to each speech segment in the audio signal to be processed is determined based on the speech keyness coefficient and tone performance coefficient of each speech segment in the audio signal to be processed.

7. The method for unstructured business data interaction based on a large language model according to claim 5, characterized in that, Principal component analysis is performed on the audio signals of each speech segment in the audio signal to be processed to determine the effective component coefficients corresponding to each speech segment in the audio signal to be processed, including: Principal component analysis is performed on the audio signals of each speech segment in the audio signal to be processed to obtain multiple principal component components of the audio signals of each speech segment in the audio signal to be processed and the characteristic values ​​corresponding to each principal component component. Feature distribution analysis is performed on the eigenvalues ​​corresponding to each principal component to determine the effective component coefficients for the corresponding speech time period.

8. The method for unstructured business data interaction based on a large language model according to claim 1, characterized in that, The audio signal to be processed is filtered according to the speech preservation level corresponding to the audio signal in each time period to obtain the target speech signal, including: Extract the speech feature sequence of the audio signal to be processed, and construct a weighted mask matrix that matches the dimension of the speech feature sequence based on the speech preservation degree of the audio signal at each time period of the audio signal to be processed. The speech feature sequence is weighted and the weighted mask matrix is ​​weighted to obtain the weighted speech feature sequence. The weighted speech feature sequence is input into the speech filtering model to obtain the target speech signal.

9. The method for unstructured business data interaction based on a large language model according to claim 1, characterized in that, Converting the target speech signal into text information includes: When the target speech signal originates from pure audio data, the target speech signal is converted into text information using an automatic speech recognition algorithm; When the target speech signal originates from audio and video data, the target speech signal is converted into first text information through an automatic speech recognition algorithm, and second text information is obtained by optical character recognition of the image frames in the audio and video data. The first text information and the second text information are then combined to form the text information.

10. A system for unstructured business data interaction based on a large language model, characterized in that, include: The signal acquisition unit is used to acquire the audio signal to be processed carrying the user's voice and at least one historical reference audio signal; The feature analysis unit is used to perform audio feature analysis on the audio signals of each time period in the audio signal to be processed based on the at least one historical reference audio signal, and to determine the speech preservation degree corresponding to the audio signal of each time period in the audio signal to be processed; the speech preservation degree is used to characterize the attention weight of the audio signal of the corresponding time period in the filtering process; The filtering processing unit is used to filter the audio signal to be processed according to the speech preservation degree corresponding to the audio signal of each time period to obtain the target speech signal; An interactive execution unit is used to convert the target speech signal into text information, input the text information into a large language model, and output feedback results for the text information.