Question and answer method and device, electronic equipment, computer readable storage medium and computer program product

By optimizing podcast audio data processing through question-and-answer methods and database construction, and generating response content with timestamps, the problem of low information density in podcasts is solved, enabling users to actively obtain multidimensional responses and efficiently retrieve information.

CN122045369APending Publication Date: 2026-05-15CHENGDU BOSS INNOVATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHENGDU BOSS INNOVATION TECH CO LTD
Filing Date
2026-02-13
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Podcast content has low information density, requiring users to invest a lot of time, and it is difficult to actively obtain targeted and multi-dimensional responses. The traditional listening mode passively limits its application scenarios and experience.

Method used

This paper provides a question-answering method that generates response content with timestamps, supports both text and audio query requests, optimizes audio data using preprocessing and speech recognition technologies, constructs a vector database for multimodal retrieval, and generates targeted response content.

Benefits of technology

It fulfills users' need to actively explore podcast content, improves information acquisition efficiency, supports multi-dimensional responses and quick location of relevant information, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045369A_ABST
    Figure CN122045369A_ABST
Patent Text Reader

Abstract

The invention provides a question and answer method and device, electronic equipment, a computer readable storage medium and a computer program product, response content corresponding to an inquiry request is generated by responding to the received inquiry request, the response content comprises response information used for replying the inquiry request and a corresponding timestamp identifier, and the timestamp identifier corresponds to the response information used for replying the inquiry request. The timestamp identification is used to locate slice information associated with the response information on the timing. According to the scheme, the inquiry request can be responded, the timestamp identifier containing the response information and used for positioning the slice information associated with the response information is generated, the requirement of a user for active exploration is met, and targeted and multi-dimensional response content is fed back.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a question-answering method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology

[0002] In recent years, digital media products containing audio, such as podcasts, have rapidly emerged as a medium for information dissemination, offering a rich and diverse range of content covering news, education, entertainment, and other fields. However, because podcasts primarily use audio, their information density is significantly lower than that of text, requiring users to invest more time in listening and resulting in lower information acquisition efficiency.

[0003] Currently, podcast content is still stuck in the traditional one-way listening model, where users can only passively listen and find it difficult to obtain targeted and multi-dimensional podcast content, which greatly limits its application scenarios and user experience. Summary of the Invention

[0004] The purpose of this invention is to provide a question-and-answer method, apparatus, electronic device, computer-readable storage medium, and computer program product to meet users' needs for proactive exploration and provide targeted and multi-dimensional response content.

[0005] In a first aspect, the present invention provides a question-and-answer method, the method comprising: In response to a received query request, generate a response content corresponding to the query request; The response content includes response information for replying to the query request, and a corresponding timestamp identifier, wherein the timestamp identifier is used to locate the slice information associated with the response information in time sequence.

[0006] Secondly, the present invention provides a question-and-answer device, the system comprising: The response module is used to respond to received query requests; A generation module is used to generate response content corresponding to the query request; The response content includes response information for replying to the query request, and a corresponding timestamp identifier, wherein the timestamp identifier is used to locate the slice information associated with the response information in time sequence.

[0007] Thirdly, the present invention provides an electronic device including a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the question-answering method described in any of the foregoing embodiments.

[0008] Fourthly, the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the question-and-answer method described in any of the foregoing embodiments.

[0009] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the question-and-answer method described in any of the foregoing embodiments.

[0010] This invention provides a question-answering method, apparatus, electronic device, computer-readable storage medium, and computer program product. In response to a received query request, it generates response content corresponding to the query request. The response content includes response information for replying to the query request and a corresponding timestamp identifier, wherein the timestamp identifier is used to locate the slice information associated with the response information in time sequence. This solution can respond to query requests and generate a timestamp identifier containing response information and the slice information associated with the response information, meeting the user's proactive exploration needs and providing targeted and multi-dimensional response content. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram illustrating an application scenario of the question-and-answer method provided in an embodiment of the present invention; Figure 2 A flowchart illustrating the question-and-answer method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the database architecture constructed in an embodiment of the present invention; Figure 4 This is one of the schematic diagrams of a visualization component for the question-and-answer method in an embodiment of the present invention; Figure 5 This is a second schematic diagram of the visualization component of the question-and-answer method in an embodiment of the present invention; Figure 6 This is a functional block diagram of the question-and-answer device provided in an embodiment of the present invention; Figure 7 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0013] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.

[0014] Audio-based digital media products have rapidly emerged as an information dissemination medium, primarily including podcasts, audiobooks, and audio programs. Taking podcasts as an example, their information density is significantly lower than text due to their reliance on voice, requiring users to invest more time and resulting in lower information acquisition efficiency. Furthermore, the unstructured nature of audio content makes podcast content difficult to retrieve and integrate efficiently, further limiting its application scenarios. To address these issues with podcast content, existing technologies have proposed several improvement measures.

[0015] For example, existing improvement measures include optimizing for noisy environments, enhancing audio quality through voice enhancement and anti-interference technologies, and improving the accuracy of speech recognition and content analysis. Additionally, measures include content moderation, such as using AI technology for compliance review to ensure platform content security. Furthermore, there are processes that combine natural language processing and multimodal analysis technologies to achieve automatic tagging, sentiment analysis, and personalized recommendations for podcast content, thereby improving content searchability and user experience.

[0016] However, most of these measures are aimed at improving the quality of audio and reviewing content, and have not improved the inherent passive listening mode of podcasts.

[0017] In view of this, embodiments of the present invention provide a question-and-answer method to improve the traditional passive listening mode into an active exploration mode, meet the user's active exploration needs, and provide targeted and multi-dimensional response content.

[0018] Please see Figure 1 This invention presents a possible application scenario for a question-and-answer system, comprising a client and a server. The client has corresponding question-and-answer application software installed, such as a digital media product containing audio. The server is the backend server of this application software. The client can receive query requests initiated by users and send them to the server. The server responds to the query request by generating a response content corresponding to the query request. Furthermore, in practical applications, the client can also execute the relevant processing steps of the server described above; this is not limited to this specific application.

[0019] This invention provides a question-answering method that can be applied to the client or server side of the aforementioned question-answering system. Figure 2 This is a flowchart illustrating a question-and-answer method provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the process includes the following steps: S11, in response to the received query request; S12, Generate the response content corresponding to the query request; The response content includes response information used to reply to the query request, as well as a corresponding timestamp identifier. The timestamp identifier is used to locate the slice information associated with the response information in time sequence.

[0020] In this embodiment, users can initiate inquiry requests based on their needs. These requests can be text-based or audio-based. The response content in this embodiment can be generated using Artificial Intelligence Generated Content (AIGC) technology, making the response content more vivid and diverse.

[0021] The question-and-answer method provided in this embodiment generates response content including response information for replying to the inquiry request, and a corresponding timestamp identifier. The timestamp identifier is used to locate the slice information associated with the response information in time sequence, so that the response information is traceable and the slice information associated with the response information can be quickly located through the corresponding timestamp identifier, thereby meeting the user's need to actively explore the digital media content.

[0022] Please refer to the following: Figure 3 The following section will first explain the database construction process.

[0023] The question-and-answer method provided in this embodiment is applied to digital media products containing audio. The query request includes text-based query requests and / or audio-based query requests for the corresponding digital media content.

[0024] The pre-built database includes a vector database mapped to the same semantic space based on text data, audio data, and metadata; a text database based on text data and metadata; and / or an audio database based on audio data. The text data, audio data, and metadata are all configured with corresponding timestamps and correspond to the same digital media content based on these timestamps. Segment information is obtained by selecting segments of the digital media content based on timestamp identifiers.

[0025] Metadata should include at least one of the following: the name of the digital media content, publisher, publication date, category tags, audio stream, notes, description, comments, citations, and resource links. Citations include references to books, papers, websites, etc.

[0026] In this embodiment, by using metadata to identify the corresponding digital media content, and configuring text data, audio data, and metadata with corresponding timestamps and assigning them to the same digital media content, precise alignment of text data and audio data can be achieved. On the one hand, this solution supports collaborative retrieval based on audio data, text data, and metadata; on the other hand, during collaborative retrieval based on audio data and text data, the corresponding text data and audio data can be accurately located, thereby achieving precise matching retrieval.

[0027] The audio data originates from the pre-processed raw audio. The pre-processing includes noise reduction, silent segment removal, speech enhancement, and / or channel conversion.

[0028] The original audio can be obtained by periodically fetching audio streams from mainstream podcast platforms or podcast subscription Uniform Resource Locators (URLs) through the Application Programming Interface (API). Alternatively, it can be obtained from local podcast files uploaded by the user. For example, it supports uploading via drag and drop on the client (supporting MP3 and WAV file formats), or direct upload via client recording (compatible with recording specifications of iOS, Android, and other systems).

[0029] Podcasts typically contain human voices, which may also include background noise, unclear pronunciation, or fluctuating volume. To address these issues, this embodiment preprocesses the obtained raw audio to resolve these problems, thereby significantly improving the accuracy of subsequent speech recognition, the effectiveness of speech information extraction, and the stability of the generation model.

[0030] Specifically, considering that podcast content may be recorded in a non-professional environment, the obtained raw audio may contain background noise such as air conditioner noise and keyboard noise. Therefore, this embodiment uses noise reduction processing to improve speech clarity.

[0031] In this embodiment, noise reduction is achieved using spectral subtraction or AI noise reduction tools. In the spectral subtraction method, the noise spectrum of the silent segment is first analyzed, and then the noise component is subtracted from the noise spectrum based on the silent segment in the original audio, thereby achieving noise reduction.

[0032] This noise reduction method is simple and intuitive in principle, with low computational complexity, small amount of computation, and extremely low hardware resource requirements. It is suitable for application scenarios with limited computing resources, such as user terminals and other devices.

[0033] In the noise reduction processing method based on AI noise reduction tools, deep learning models with noise reduction functions (such as RNNoise (an open source audio noise reduction library), Adobe Enhance Speech (a speech enhancement and noise reduction tool)) are used to separate human voices and noise in the original audio, thereby achieving noise reduction.

[0034] This noise reduction method has the advantages of excellent noise reduction effect and high adaptability. It can handle complex and variable noise and can automatically identify and adapt to different noise environments and acoustic conditions.

[0035] Furthermore, considering that podcasts may contain long pauses or silent segments, removing these segments can reduce the amount of data and improve subsequent processing efficiency. Therefore, in this embodiment, the original audio is processed by removing silent segments.

[0036] It can detect silent segments in the original audio based on energy thresholds and combine them with a speech activity detection (VAD) algorithm to remove silent segments.

[0037] To avoid a sense of speech interruption and preserve natural pauses in speech, in this embodiment, among the detected silent segments, silent segments with a duration less than a set duration (e.g., 0.5 seconds) are retained, so that the final processed speech data has natural pauses and is more in line with the effect of real speech.

[0038] In addition, considering that the original audio may have problems such as volume fluctuations and echoes, this embodiment can also perform speech enhancement processing on the original audio.

[0039] In this embodiment, the original audio undergoes equalization processing to adjust frequency bands and make vocals clearer, such as enhancing mid-frequency frequencies and reducing low-frequency noise. The original audio may also have issues like clipping or distortion, which need to be corrected. Therefore, for segments in the original audio such as whispers and sudden laughter, the volume fluctuations are balanced. Dynamic range compression can balance volume fluctuations, making the speech smoother and thus avoiding frequent volume adjustments by the listener.

[0040] If there is an echo in the recording environment, the sound will sound muffled. To address this, this embodiment uses a dereverberation model based on a convolutional neural network or spectral subtraction combined with reverberation time estimation to remove the echo, thereby achieving the purpose of echo elimination or reverberation reduction and improving sound quality.

[0041] In the original audio, vocals may be mixed with background music and sound effects. In order to extract clean vocals, this embodiment uses a separation model to separate the original audio. The separation model can be a model such as Demucs (Facebook open source music source separation model), which can separate drums, bass and vocals from accompaniment.

[0042] If the podcast is in stereo, the original audio can also undergo channel conversion. Since speech is usually in the middle channel, converting the original audio to mono reduces the amount of data required for subsequent processing. In this embodiment, a standardized sampling rate and bit depth can be used to ensure consistency and avoid problems in subsequent processing.

[0043] The original audio may also contain sounds such as clicking or popping, and there may also be unclear pronunciation. Based on this, this embodiment can also eliminate sounds such as clicking and popping in the original audio, and adjust unclear pronunciation manually or automatically to improve the clarity of speech.

[0044] Speech recognition and speech information extraction are performed on the audio data obtained after the above preprocessing.

[0045] Speech recognition can be implemented in various ways, such as using common open-source models. Some common open-source models are as follows: FunASR is a basic speech recognition toolkit that provides a variety of functions, including speech recognition (ASR), speech endpoint detection, punctuation recovery, language modeling, speaker verification, speaker separation, and multi-person dialogue speech recognition. FunASR provides convenient scripts and tutorials, supports inference and fine-tuning of pre-trained models, and allows the use of different models to extract corresponding information.

[0046] SenseVoiceSmall: An open-source multilingual speech recognition model with various speech understanding capabilities, covering automatic speech recognition, language recognition (LID), emotion recognition (SER), and audio event detection (AED).

[0047] paraformer-zh: A non-autoregressive end-to-end Chinese speech recognition model that can perform speech recognition with timestamp output, and also has versions based on long text and speaker logs and extraction.

[0048] Qwen-Audio: A large-scale multimodal audio-text model. Qwen-Audio not only transcribes input audio, but also has deeper capabilities such as semantic understanding, sentiment analysis, audio event detection, and voice chat.

[0049] cam++: A technique for visualizing the decision-making processes of convolutional neural networks (CNNs), supporting speaker identification / segmentation.

[0050] MFCCA: A multifractal cross-correlation analysis technique. The goal of multi-talker speech recognition (Multi-talker ASR) is to recognize speech containing multiple speakers and to correctly identify speech with highly challenging speaker overlap.

[0051] SOND: Speaker Log: Given raw speech / voiceprint information of several speakers, identify and track these speakers in speech segments.

[0052] In this embodiment, speech recognition is performed based on the current open-source model. This allows for speech recognition with a low implementation cost, flexible customization of the model based on requirements, and supports independent deployment of the model, ensuring the privacy of sensitive data.

[0053] In addition, commercial API speech recognition solutions, such as the Volcano Engine solution and the iFlytek solution, can also be used to perform speech recognition. The Volcano Engine solution supports transcribing audio files (≤4 hours) into text data, and has built-in functions such as automatic punctuation, semantic smoothing, number normalization, and intelligent sentence segmentation, which can be combined as needed. The iFlytek solution's speech transcription interface can convert audio files (within 5 hours) into text data through integrated development, returning word segmentation form, complete sentence form, word and sentence information, word and sentence timestamps, word attributes, multiple candidate words, intelligent grammatical format conversion, and multiple speaker separation, etc.

[0054] In this embodiment, using commercial API speech recognition can reduce the cost of maintaining infrastructure and model operation and maintenance, and can achieve a higher recognition accuracy.

[0055] In addition to the above-mentioned methods for speech recognition, speech information can also be extracted by combining the recognition results with podcast metadata information and a large speech-text dialogue model. Speech information extraction may include the following operations: Hierarchical speech recognition: Continuous audio is segmented using a speech activity detection method, and roles are labeled using a speaker separation model. The output is a sentence-by-sentence text with timestamps and speaker labels.

[0056] Semantic unit segmentation: Based on semantic similarity clustering, continuous sentences are aggregated into paragraphs, and chapter boundaries are detected by combining prosodic features to generate a three-level summary (chapter / full text / speaker).

[0057] Key information extraction: Use a joint entity recognition model to extract keywords (people, places, events), to-do items (such as "need to read XX paper"), and scene tags (interviews, debates, tutorials).

[0058] Cross-modal association construction: establishing "audio segments" Text transcription The "External Resources" citation network automatically parses hyperlinks in notes and associates them with the titles of books and papers mentioned in the program with the corresponding audio locations.

[0059] The final audio information is as follows: a. Division: Divide the data into chapters based on the recognition results, which serve as the basis for contextual search.

[0060] b. Summary: Summary of each chapter, the entire text, and each speaker's remarks, as well as a review of key questions and answers, etc.

[0061] c. Key points: keywords, to-do items, key content, and scenario identification.

[0062] The divisions, summaries, and key points extracted above are saved to the database as extended fields after subsequent processing.

[0063] Based on the above, after optimization processing, the audio data is encoded into relatively aligned text vectors and audio vectors by a text-audio dual encoder. The text vectors are stored in a text database, and the audio vectors are stored in a vector database. The optimization processing includes: timbre adaptation optimization processing, long-term context modeling processing, and / or scene-aware feature enhancement processing.

[0064] In this embodiment, the Wav2Clip model is used as the basic architecture, whose dual encoder structure (audio encoder + text encoder) naturally supports cross-modal alignment. The following optimizations can be performed on the audio data: Timbre-adaptive optimization: Based on the LibriSpeech speech corpus, historical audio (including multi-person dialogues and background music) is introduced for model fine-tuning to enhance the model's robustness to overlapping speech and non-stationary noise. The fine-tuned model is then used to perform timbre-adaptive optimization on the audio data.

[0065] In this embodiment, the model is fine-tuned based on historical audio, enabling the fine-tuned model to accurately capture the voice fingerprint of the target person or scene, including information such as timbre and intonation. This allows the model to be used to optimize the audio data to be processed, including its timbre.

[0066] Long-term context modeling processing: Expand the short-duration (e.g., 3 seconds) input window of the original model to a long-duration (e.g., 10 seconds) input window, cover the complete sentences in the audio data with a sliding window (e.g., step size of 5 seconds), and use the Transformer layer in the model to capture long-range dependencies.

[0067] In this embodiment, extending the existing short-duration processing window to a long-duration window can improve the model's ability to process long audio files and better preserve contextual information. In this implementation, an overlapping area needs to be set between the two sliding windows to ensure smooth segmentation and avoid abrupt transitions.

[0068] Scene-aware feature enhancement processing: A scene classification branch (identifying 12 types of scenes such as interviews, speeches, and debates) is connected in parallel to the output layer of the model. Scene labels are embedded into vectors (for example, 16-dimensional scene encoding is added to the original 128-dimensional vector, resulting in a 144-dimensional vector). Then, the following processing flow is performed: Feature extraction: 80-dimensional Mel spectrograms are extracted from each frame of audio data, and the spectrograms are concatenated into an 800×80 matrix in 10-second windows and input into the Wav2Clip model.

[0069] Vector generation: After the Wav2Clip model outputs a 144-dimensional vector, it is reduced to 128 dimensions by PCA to reduce the index storage pressure.

[0070] In this embodiment, scene labels are introduced into the Wav2Clip model, enabling the model to perform targeted analysis and processing based on the audio characteristics of the corresponding scene when processing audio data in different scenes, thereby improving the audio data processing effect in different scenes.

[0071] In addition, the text data in the database includes structured fields, unstructured data, and reference relationships corresponding to digital media content. After being vectorized by a text encoder, the text data is stored in a vector database in the form of sparse text vectors and dense text vectors.

[0072] In this embodiment, to address the multi-granularity retrieval needs of podcast text, BGE-m3 (a multi-functional text embedding model) is used as the core text encoder to achieve the joint generation of sparse and dense vectors, combining the advantages of traditional keyword matching and deep semantic retrieval. In this embodiment, OpenSearch or ElasticSearch are uniformly used for storage, with the text semantic fields stored in the vector database `text`. An ik segmentation toolkit (a lightweight Chinese word segmentation toolkit) is added, and BM25 retrieval is implemented by default.

[0073] The text data that needs to be vectorized includes digital media content and audio speech recognition results.

[0074] Fields requiring structured processing in digital media content include program name, publisher, release date, and category tags. Unstructured data in digital media content includes audio streams, program descriptions, notes, and listener comments. Citations in digital media content include external resources mentioned in the program (books, papers, websites). The vectorized structured data, unstructured data, and citations in digital media content are stored in a vector database according to their mapping relationships with corresponding metadata. Additionally, the raw digital media content in the text data is stored in a text database according to its mapping relationship with corresponding metadata.

[0075] For the audio speech recognition results in the text data, the following vectorization processing is performed in conjunction with the corresponding metadata: Segmentation: Sentence-by-sentence speech recognition results, timestamp speaker + text; Summary: This includes summaries of chapters, the entire text, individual speakers' remarks, and a review of key questions and answers; Key points: keywords, to-do items, key content, and scenario identification.

[0076] After all the sparse and dense vectors of the above text data have undergone vectorization processing and are stored in the vector database, a hierarchical index structure is established in the vector database as shown below: Main index: Stores program-level metadata (name, category tags, publisher) and its statistical characteristics (listening popularity, update frequency); Segment Index: Stores audio vectors, text vectors, keywords, scene tags, and associated external resources by time window (default 5 minutes); Abstract Index: Stores compressed vectors of chapter summaries, speaker viewpoint summaries, and Q&A reviews for quick location of macro topics.

[0077] The above methods allow for the construction of a complete database, including a vector database, a text database, and / or an audio database. Furthermore, the three-level hierarchical indexing architecture helps improve the efficiency and accuracy of multimodal retrieval. Introducing cross-modal vector alignment technology maps audio and text to a unified semantic space based on metadata, enabling efficient collaborative retrieval of audio, text, and metadata. In this embodiment, through multi-level vectorization, heterogeneous data association, and adaptive index optimization, podcast audio, text, metadata, and external resources are integrated into a unified semantic network, supporting millisecond-level cross-modal retrieval.

[0078] Based on the above, in practical applications, when faced with a user's query request, a search can be performed in the database to generate the corresponding response. Specifically, the step of generating a response to a received query request can be implemented in the following way: The query request is used to search the relevant database to determine the target search results, which correspond to the original slice content. Response content is then generated based on the target search results.

[0079] Here, the original slice content refers to excerpts of text data, audio data, etc., stored in the database. Based on the query request, the database is searched to obtain target retrieval results corresponding to the original slice content, and then response content is generated based on the target retrieval results.

[0080] To provide faster and better responses to user inquiries, this embodiment analyzes the user's inquiry intent and then performs targeted searches based on that intent. Therefore, in this embodiment, the step of searching the relevant database based on the inquiry request to determine the target search results can be implemented in the following way: Inquiry requests are analyzed to determine the corresponding inquiry intent, and a corresponding retrieval strategy is determined based on the inquiry intent; the database is then searched according to the determined retrieval strategy to determine the target retrieval results.

[0081] In this embodiment, the query request can be input into the corresponding intent recognition model to determine the intent type of the query intent, and the corresponding intent parameters can be extracted according to the intent type.

[0082] Then, based on the intent type and intent parameters, a retrieval strategy is determined. The retrieval strategy includes a retrieval pattern that matches the query intent and a result reordering weight corresponding to the retrieval pattern. The retrieval patterns include a text retrieval pattern corresponding to factual query requests, a vector retrieval pattern corresponding to opinion-based query requests, and / or a graph relationship retrieval pattern corresponding to resource request query requests.

[0083] In this embodiment, the intent recognition model can analyze the inquiry intent of the request from multiple dimensions, such as fact query, opinion summary, resource acquisition, scenario positioning, to-do list, comparative analysis, and open discussion. Based on the analysis results, the inquiry intent is then categorized into one of three intent types: fact query, opinion summary, and resource acquisition.

[0084] In this embodiment, a dedicated parser can be deployed for each type of intent, and the corresponding intent parameters can be extracted according to the intent type. For example, for the fact query type, the entity corresponding to the fact can be extracted.

[0085] The retrieval strategy is determined based on the intent type and intent parameters. The retrieval strategy includes retrieval mode and result reordering weights. Among the retrieval modes, the text retrieval mode corresponding to factual query requests mainly performs retrieval based on precise fragment matching. In this retrieval mode, time importance is relatively low. Therefore, the time decay weight in its corresponding result reordering weight can be set to a small value. For example, the time decay weight can be multiplied by a number less than 1, such as 0.3.

[0086] The vector retrieval mode corresponding to opinion-based query requests is mainly based on the chapter summary priority query method. In this retrieval mode, the speaker's authority is relatively important. Therefore, in the corresponding result reordering weight, the speaker authority weight can be set to a large value. For example, the speaker authority weight can be multiplied by a number greater than 1, such as 1.2.

[0087] The graph relationship retrieval mode corresponding to resource request type query requests mainly retrieves information by traversing the association graph constructed by external reference relationships. In this retrieval mode, the publication time of the referenced content is relatively less important. Therefore, in the corresponding result reordering weight, the weight of the publication time of the referenced content can be multiplied by a number less than 1, such as 0.8.

[0088] In this embodiment, a multi-path retrieval method is adopted in the process of searching the database according to the determined retrieval strategy to determine the target retrieval results.

[0089] As one possible implementation, when the user-initiated query request is a text-based query request, the step of searching the database according to the determined retrieval strategy to determine the target retrieval results can be implemented in the following way: In response to a text-based query request, the text-based query request is converted into sparse vectors and dense vectors, and vector retrieval is performed in a vector database based on the sparse vectors and dense vectors to determine the corresponding text vector retrieval results; and text retrieval is performed in a text database based on the text-based query request to determine the corresponding text retrieval results.

[0090] As mentioned above, databases include vector databases, text databases, and audio databases. Text data, audio data, and metadata in a vector database are mapped to the same semantic space.

[0091] When a user initiates a text-based query request, it can be processed in two ways: firstly, the text query request can be vectorized, including sparse and dense vectors, and then retrieved from a vector database based on these vectors; secondly, it can be retrieved directly from a text database based on the text query request itself.

[0092] For text-based query requests, keyword extraction and context completion can be performed.

[0093] During the retrieval process, for short texts in the database, such as program names and category tags, a word frequency (TF) statistical model can be used to perform word frequency statistics, and corresponding weights can be assigned to each short text based on the TF statistical results. Furthermore, for non-semantic fields such as category tags and scene tags, TF-IDF can be used to assign corresponding weights to the texts corresponding to the category tags and scene tags.

[0094] In vector retrieval, the vector engine can be used for nearest neighbor search and supports multi-field joint search, such as simultaneously matching a user's textual query request with ASR text and chapter summaries.

[0095] In this embodiment, a hybrid retrieval of sparse and dense vectors is performed, returning the search results ranked in the top preset positions by search score, such as the top 200 results. The search scores are then normalized to a set range, such as [0,1]. The search score is obtained by weighting the sparse vector search results and the dense vector search results, as shown below:

[0096] in, The retrieval score of the i-th retrieval result is a comprehensive score combining keyword matching and semantic matching; a and b represent the weights of the sparse vector retrieval result and the dense vector retrieval result, respectively. The BM25 score, obtained based on sparse vectors, represents the degree of matching between the i-th search result and the literal / keyword of the query term in the textual query request. The similarity between vectors is usually expressed as the cosine similarity after embedding, representing the similarity of the first vector. The semantic similarity between a search result and a text-based query request.

[0097] Furthermore, when the user-initiated query request is an audio query request, as a possible implementation, the step of searching the database according to the determined retrieval strategy to determine the target retrieval results can be implemented in the following way: In response to an audio query request, the audio query request is converted into an audio vector, and a vector retrieval is performed in a vector database based on the audio vector to determine the corresponding audio vector retrieval result; speech recognition is performed on the audio query request to convert it into a corresponding text retrieval request, and a text retrieval is performed in a text database based on the text retrieval request to determine the corresponding translation retrieval result; and / or, an audio retrieval is performed in an audio database based on the audio query request to determine the corresponding audio retrieval result.

[0098] In this embodiment, for audio query requests, the first aspect is to convert the audio query request into an audio vector, and directly compare the converted audio vector with the audio vectors in the vector database to obtain the audio vector retrieval result. In this embodiment, the audio query request can be the original audio or a humming excerpt, suitable for retrieval of music podcasts, etc. The retrieval score can be based on the CLAP (Contrastive Language-Audio Pre-trained Model) audio vector similarity, returning the audio vector retrieval results ranked in the top preset positions, such as the top 150, and the retrieval score can be normalized to a set range.

[0099] The second aspect involves converting audio query requests into text retrieval requests, and then performing searches in a text database based on these text retrieval requests. The third aspect involves directly searching an audio database based on the audio query requests.

[0100] Retrieve from at least one of the three aspects mentioned above for audio query requests.

[0101] To achieve efficient fusion of text and audio dual-modal retrieval results, this embodiment performs fusion processing on the retrieval results obtained from the text retrieval channel and the audio retrieval channel. Based on this, the step of searching the database according to the determined retrieval strategy to determine the target retrieval results may further include the following steps: The search results are fused from text vector search results, text search results, audio vector search results, translated search results, and / or audio search results to obtain corresponding fused search results; the target search result is determined from the fused search results based on the result re-ranking weight.

[0102] In this embodiment, in the step of obtaining the fused search results, the timestamps of the corresponding text data, audio data, and metadata can be used to perform timestamp alignment and deduplication on the text vector search results, text search results, audio vector search results, translated search results, and / or audio search results to obtain the corresponding fused search results.

[0103] In this embodiment, due to the use of a multi-path retrieval method, after timestamp alignment of the retrieval results, the retrieval results hit by different retrieval methods may have time overlap. For example, text retrieval result A hits the content corresponding to the podcast from 10:00 to 12:00, while audio retrieval result B hits the content corresponding to the podcast from 11:30 to 13:00. The overlap in time exceeds a preset threshold (e.g., 50%). In this embodiment, deduplication can be performed, merging text retrieval result A and audio retrieval result B into a supersegment (i.e., the content corresponding to the podcast from 10:00 to 13:00), and using the higher retrieval score between text retrieval result A and audio retrieval result B as the retrieval score of the supersegment, which will then participate in the subsequent retrieval result ranking.

[0104] Furthermore, if there is a conflict in the speaker labels of the retrieved text channel results (including text vector retrieval results and / or text retrieval results) and audio channel results (including audio vector retrieval results, transcribing retrieval results and / or audio retrieval results), a penalty weight of less than 1 can be applied after the two are merged.

[0105] After obtaining the fusion search results through the above methods, the target search results are determined from the fusion search results based on the result reordering weights. Specifically, a non-linear weighting mechanism can be used to reorder the fusion search results based on the result reordering weights to obtain the corresponding reordered search results.

[0106] In this embodiment, the text channel results and audio channel results in the fusion search results are sorted from high to low according to their search scores. The sorted text channel results and audio channel results are stored in the constructed global candidate pool, and the RRF score for each result is calculated according to the following formula:

[0107] The Reciprocal Rank Fusion Score (RRF Score) is the final comprehensive score used to rank the global candidate pool.

[0108] Text channel weights are used to control the importance of text channel results in the final ranking.

[0109] Audio channel weights are used to control the importance of audio channel results in the final ranking.

[0110] The smoothing constant, usually set to 60, is used to prevent audio or text data ranked very high (such as number 1) from dominating the results, thus giving audio or text data ranked lower a chance.

[0111] : The ranking of the text channel results, the position of a data point in the text channel results (e.g., if it is the 1st place, the rank is 1).

[0112] The ranking of the audio channel results is similar to the position of the data within the audio channel results.

[0113] Thus, following the above method, the fused search results can be re-ranked based on RRF scores to obtain re-ranked search results. On this basis, the slice content corresponding to the re-ranked search results is expanded with related content, and then subjected to secondary re-ranking and secondary deduplication to obtain the target search results.

[0114] In this embodiment, the slice content corresponding to the result with the highest RRF score can be selected based on the reordering retrieval results.

[0115] Perform related content expansion on the selected slice content, such as expanding sentence-level results into the context to the entire paragraph, to ensure semantic integrity.

[0116] The sliced ​​content after related content expansion can undergo secondary reordering, such as multi-feature fusion reordering, which can be achieved by pre-training a BERT (Transformer-based bidirectional encoder) reordering model. The BERT reordering model can be trained by inputting the following features: The semantic similarity between the slice content corresponding to the reordered search results and the query request; The popularity rating of the chapter containing the sliced ​​content corresponding to the reordered search results; The speaker authority corresponding to the slice content of the reordered search results (based on the publisher influence model). The time decay factor of the slice content corresponding to the reordered search results (new content takes precedence).

[0117] Based on this, a second deduplication process is performed on the slice content after the second reordering. The strategy for the second deduplication process is to merge slice content with overlapping timestamps and retain the merged slice content with the highest confidence. Finally, the slice content with the highest RRF score in the first preset position is returned as the final target retrieval result.

[0118] Based on the target retrieval results obtained through the above methods, response content is generated. In this embodiment, the response content can be generated by a generative large language model. Specifically, the steps for generating response content based on the target retrieval results can be implemented in the following ways: The search enhancement generation technique is used to inject corresponding prompt words into the target search results; the prompt words are then input into a generative large language model to generate response content.

[0119] In this embodiment, the prompt word instruction includes the following information: Inquiry requests and target retrieval results; instructions for setting roles for generative large language models; instructions for selecting criteria for target retrieval results; instructions for output requirements for response content; instructions for reasoning constraints on response content.

[0120] In this embodiment, the retrieved target search results are input into a generative large language model (such as GPT-4 or deepseek-v3). A retrieval augmentation generation (RAG) framework is used, injecting the query request as a prefix into the generative large language model. Constraint control and source tracing verification are employed to improve the quality and credibility of the generated answers.

[0121] In this embodiment, the output requirement instructions for the response content included in the prompt word instructions are mainly controlled by triple constraints. These constraints aim to limit and guide the generation process of the large language model from different dimensions, ensuring that the large language model can accurately utilize the search results and generate response content that meets expectations. The output requirement instructions mainly include the following: 1. Evidence weight constraint: guides the model to make judgments based on the reliability of the search results.

[0122] Priority: When integrating information, the constraint model must prioritize the slice content with the highest retrieval score (e.g., RRF Score ≥ 0.8) as the core argument.

[0123] Consistency: When there are contradictions among the retrieved slices, the model must select the most credible slice based on its weights and is constrained not to simultaneously reference contradictory low-weight slices. The Prompt requires the model to declare in its answer which piece of evidence it has selected and to ignore conflicting evidence with lower weights.

[0124] 2. Format and content constraints: Ensure that the structure, style and essential elements of the generated response content meet the system requirements.

[0125] Structured output: Forces the model to output the response content in a predefined format (such as JSON or Markdown bullet points) so that the system can perform subsequent parsing and footnote marking.

[0126] Source tagging: Source tagging forces the model to insert a placeholder or temporary reference ID at the end of each sentence that cites an opinion or fact, so that the system can replace it with the final footnote tag. The prompt instruction is as follows: "Insert a placeholder {cite: chunk_id} at the end of each quoted sentence." Spoken to Written Language Conversion: During generation, the constraint model must clean and transform colloquialisms, redundancies, and interjections in the content slices to ensure that the output response is professional and fluent written language. A prompt instruction could be: "Please remove redundancies and interjections from the context and convert the spoken language into fluent written language." 3. Negative constraints: to prevent the model from creating illusions.

[0127] Knowledge Boundaries: Models are strictly prohibited from using their own training knowledge or any unretrieved context to answer questions. The primary rule for system prompts is: "You are a pure knowledge synthesis engine and can only use..."<retrieved_context> The information is in the text. If you cannot answer, you must reply with a pre-defined rejection statement, such as 'Unable to answer based on the context'.

[0128] The reasoning restriction instructions in the prompt words are mainly used to constrain the model to only perform lightweight reasoning and summarizing of the retrieved slice content, and not to perform in-depth extrapolation or subjective evaluation, so as to avoid over-interpretation.

[0129] The following example illustrates prompt word instructions input into a generative large language model: #Identity and Responsibilities You are based on podcast content Advanced knowledge synthesis engine Your sole task is to answer user queries strictly according to the provided JSON-formatted retrieval context. Your core responsibility is: Accurate, traceable, and structured Integrate local information.

[0130] ## Follow the rules ### I. Refer to the usage rules Core Priority Principles: Evidence fragments with the highest `relevance_score` must be used first.

[0131] Conflict resolution mechanisms: If contradictions or inconsistencies are found among the evidence, you are strictly limited to choosing only one option. The highest relevance score Low-scoring segments that are used as the sole argument and ignore all conflicting points.

[0132] Citation threshold: Only `relevance_score` is allowed. Greater than 0.6 Substantive citations of fragments.

[0133] ### II. Format and Content Language conversion (spoken language cleansing): You must remove all colloquial expressions (such as interjections, repetitions, and redundancy) from the search text and transform the content into... Fluent, professional, and concise written language .

[0134] Information Summary: A comprehensive answer must be provided, taking into account information from all qualified sources (text transcription, audio summarization, etc.).

[0135] Source tracing marker (mandatory verification prefix): For every cited opinion or fact in the answer, a quotation mark must be inserted at the end of the sentence. Temporary reference marker The format is `[CID:`<chunk_id> This ID must correspond exactly to the `chunk_id` of the original retrieved fragment.

[0136] Content requirements: The sentence you generate must be Direct mapping or light summary of the original fragment This is to ensure that subsequent verifications are passed.

[0137] ### III. Negative Constraints Knowledge Boundaries (Anti-Hallucination): This is the strictest rule. The use of any self-trained knowledge or information not provided by the search context is strictly prohibited.

[0138] Refusing to answer: If the retrieved context does not meet the minimum evidence requirements for an answer (e.g., all relevant scores are below a threshold), you must reply: "Based on the podcast content currently retrieved, I cannot provide a complete answer."

[0139] ## Final Output Format You must strictly follow the following structured format to output your final answer.

[0140] Key conclusions A one-sentence summary of the answer to the user's query.

[0141] Detailed analysis Using Markdown breakpoints ( (This section elaborates on specific details and citations.)

[0142] In this embodiment, for both audio and text-based query requests, multi-path retrieval results are aligned, fused using RRF (Regressive Language Ranking), and reordered using a non-linear weighting mechanism to achieve adaptive optimization and integration of cross-modal retrieval results. The final retrieval results are transformed into structured prompt prefixes, and prompt word instructions with constraints guide the generative large language model to generate response content. The generative large language model can quickly respond to user query requests, generating and providing feedback on response content in real time, achieving an instant listening experience. Furthermore, through the synergistic innovation of non-linear RRF fusion and generative control technology, this solution achieves a breakthrough improvement in accuracy, credibility, and interpretability in podcast question-and-answer scenarios, providing a solution for the knowledge-based application of multimodal content.

[0143] This solution pre-creates vector, text, and audio databases and aligns them based on timestamps. It supports various user query needs, including searching text by audio and searching audio by text. Compared to traditional simple text or audio search methods, it can meet users' diverse search needs and greatly improve the user experience.

[0144] Based on this, in this embodiment, the step of generating response content based on the target retrieval results further includes: Perform content validation and / or source tracing validation on the generated response content, and output the response content after the validation passes.

[0145] In this embodiment, the source verification of the response content is a separate step, usually performed after the model generates the response content and before returning it to the user. The verification is mainly performed from the following aspects: 1. Citation Parsing: The system parses the response content output by the model and extracts all temporary citation markers (e.g., `{cite: chunk_1}` or...). ).

[0146] 2. Extracting Statements: For each referenced statement, the system extracts that statement.

[0147] 3. Reference anchoring: Core verification: The system performs a semantic similarity comparison between the referenced statement and the corresponding original slice content.

[0148] Pass condition: A citation is considered valid only if the similarity score (such as cosine similarity) between the cited statement and the original slice content is higher than a preset threshold (e.g., 0.90).

[0149] 4. Corrections or warnings: Failure handling: If the verification fails (the similarity is too low, meaning the model generated an illusion or over-interpreted), the system has the following two handling methods: Correction: Try replacing the statement with an expression that is more similar to the original slice content (achieved by calling another correction model).

[0150] Downgrade: Remove the footnote and mark the response as "lacking strong supporting evidence".

[0151] Successful processing: The temporary citation marker was replaced with the final clickable footnote marker (e.g., ...). ).

[0152] Finally, in terms of user interaction, the product implements clickable footnote markers (such as clickable footnote markers). The steps to pop up the context of the original slice content are as follows: 1. Front-end binding: The front-end page marks the footer. Bind to a click event.

[0153] 2. Data transmission: After the click event is triggered, the front end obtains the `chunk_id`, `episode_title`, and `timestamp` corresponding to the footnote.

[0154] 3. Interface Display: A pop-up sidebar or floating window displays the following information: Original slice content: Complete search text content (highlighting the core sentences quoted in the answer).

[0155] Metadata: podcast name, episode number, speaker, etc.

[0156] Jump Link: Provides a clickable link or play button that allows users to jump directly to the corresponding timestamp in the podcast app or web player to listen.

[0157] The above methods achieve a complete closed-loop tracing experience from answer to evidence to original audio. This ensures the integrity of the evidence chain by directly verifying the source of the answer and tracing back to the original audio, ensuring that the information is not taken out of context or generated out of thin air. Verifying the original audio avoids semantic biases introduced during translation, enhancing the credibility of the conclusions.

[0158] like Figure 4 As shown, the generated response content can be displayed on the interface, and the response content includes footnote markers (such as...). [1] Clicking the footnote marker in the response content will bring up the source tracing viewer, which displays the original slice content and timestamp. The original slice content may include the title of the original slice content (e.g., ...). Figure 4The "Future of AI Startups (Vol.42)" in the article), and the timestamp information of the original slice content (such as... Figure 4 The audio data includes segments from 08:15 to 09:30, the original slice content, and the corresponding text transcription. The playback timeline of the audio data can be displayed in the source viewer.

[0159] This solution provides a source tracing function, making the source of the response content transparent. Users can directly view relevant evidence supporting the response (such as key paragraphs, summaries, etc.). Furthermore, based on timestamps, it can directly locate the associated original segment content, quickly focusing on the core content. By tracing back to the original segment content, such as in the implementation of audio playback of the original segment content, implicit information that cannot be conveyed by text (such as the speaker's emotions, emphasis, etc.) can be captured by listening to the tone of voice and background sounds in the audio (such as applause, pauses), allowing for intuitive and accurate acquisition of effective information. In this embodiment, in addition to adding footnotes to the response content and jumping to the original segment content based on the footnotes, such as... Figure 5 As shown, the response content may include specific response information and a corresponding timestamp. The timestamp identifies the slice information associated with the response information in time sequence. This slice information includes the original slice content corresponding to the query request. Based on this, the question-and-answer method provided in this embodiment may further include the following steps: In response to the first trigger command for the timestamp identifier, the original slice content is fed back.

[0160] In this embodiment, the first trigger command for the timestamp identifier can be anything, such as a mouse click or double-click. When the corresponding timestamp identifier is triggered, the original slice content corresponding to that timestamp identifier can be fed back. The form of feedback can be either through voice playback or by converting the original voice into text, etc.

[0161] like Figure 5 As shown, the response content includes specific response information, which is marked with a timestamp (e.g., @05:10). The timestamp identifies the original slice content (e.g., Deepin Technology) associated with the response information in terms of time sequence. When the timestamp in the response information (e.g., @05:10) is triggered, the corresponding original slice content (e.g., the specific slice content in Deepin Technology) can be fed back. The feedback of the original slice content can be in the form of audio data playback, or it can be displayed on the interface as its corresponding text data or translated text data.

[0162] In addition to the original slice content corresponding to the query request, the slice information may also include contextual content, summary content, citation content, and / or explanatory content corresponding to the original slice content. Based on this, the question-answering method provided in this embodiment may further include the following steps: In response to a second trigger command for a timestamp identifier, it provides contextual content, summary content, citation content, and / or explanatory content.

[0163] Similarly, the second trigger command could be a mouse click or double-click on a timestamp identifier. When the timestamp identifier is triggered, it can also provide contextual information, summary content, citation content, and / or explanatory content of the original slice. For example, a relational graph can be constructed based on the citation content to visually display the citation content. Figure 5 As shown in the diagram, the citation relationships between the original slices are visually displayed in the cross-podcast association graph. For example, there is a citation relationship between "digital ethics" and "future health theory", "deep technology" and "digital ethics", and "future health theory" and "AI will society" have a common discussion relationship.

[0164] In this embodiment, by providing feedback on the citation content of the original slice content, the citation relationships and common discussion relationships of the original slice content are intuitively displayed. By following the citation paths of the citation relationships, users can understand the most original slice content that provides the core content (slice content nodes that are cited extensively). Furthermore, through common discussion relationships, seemingly unrelated original slice content that shares common discussion points can be automatically identified, triggering cross-disciplinary insights and helping users to comprehensively understand the connections between different fields.

[0165] The summary content may include a summary of the original slice content, and may also include a summary relating the cited content to the original slice content. The explanatory content mainly serves to explain and clarify, such as the reasons for recommending the content. Based on this, the response content also includes a comparative analysis of the query request and the original slice content. The question-and-answer method provided in this embodiment may further include the following steps: In response to a third trigger command targeting a timestamp identifier, provide feedback and comparative analysis results.

[0166] In this embodiment, for example, a user can upload an audio clip (such as a meeting recording clip), and the system will automatically extract the audio vector of the clip and search for similar podcast segments to generate comparative analysis content ("The viewpoint you provided is 78% similar to Professor Li's argument in the XX podcast").

[0167] This solution supports comparative analysis of arguments between audio segments, automatically identifying the core arguments, evidence, and conclusions in two audio clips, and visually displaying similarities and differences through methods such as comparison tables or visualization graphs. This approach directly provides users with similar or contradictory audio content, helping them quickly and comprehensively understand various viewpoints in the relevant field. In this embodiment, the response content also includes a timeline and derived segments arranged sequentially along the timeline. The derived segments are associated with the original segments, and the timestamp is configured as a movable control on the timeline. The response content provided in this embodiment also includes the following steps: In response to the timestamp being moved to the position corresponding to the derived slice content, the derived slice content is fed back.

[0168] like Figure 5 As shown, on the visualization interface, the various derivative slices are arranged sequentially on the timeline, and each derivative slice has a corresponding timestamp. When the corresponding timestamp on the timeline is triggered, the corresponding derivative slice content will be displayed, such as displaying the corresponding text or playing the corresponding audio.

[0169] It should be noted that in this embodiment, the original slice content can also be sorted according to the timeline (e.g., Figure 5 The lower left area, within the audio playback controls and event timeline, displays an overview of data privacy risks and ethical issues. The timestamp corresponding to the original slice content is configured as a movable control on the timeline (e.g., 05:10, 14:25). Similarly, when the timestamp of the original slice content on the timeline is triggered, the corresponding original slice content is displayed.

[0170] In this embodiment, by arranging the derived slice content and the original slice content in sequence on the timeline, the timeline is used to visually display the results of the contextual analysis. This can visualize the abstract temporal relationships between the various derived slice content and the original slice content. Furthermore, by combining timestamp identifiers and triggering jumps based on timestamp identifiers, dynamic correlation between content is achieved.

[0171] In this embodiment, after the step of responding to the timestamp identifier being moved to the position corresponding to the derived slice content and feeding back the derived slice content, the method further includes the following steps: In response to the fourth trigger command for the timestamp identifier, feedback is provided regarding the evidence information for the derived slice content.

[0172] In this embodiment, the derived slice content is associated with the original slice content. This can be understood as recommending derived slice content based on certain reasons, determined from the original slice content. For example... Figure 5The recommended podcast segments shown include derivative segments related to the original segment (such as "Cybersecurity Frontiers" and "Tech Observer"). To improve interpretability, in addition to providing feedback on the derivative segment content, it can also provide evidence related to that derivative segment content, such as high similarity between the derivative and original segments, and strong thematic relevance. For example, content similarity and thematic relevance can be quantified as follows: Figure 5 Values ​​like [8.8] and [8.5] ​​indicate that the higher the value, the stronger the content similarity and thematic relevance.

[0173] In summary, the question-answering method provided in this embodiment deeply adapts retrieval enhancement and generation technology to digital media products, constructing a closed-loop system encompassing multimodal retrieval, dynamic enhancement, and reliable generation. Specifically, this includes the enhancement and fusion processing of multimodal data (audio, text, and metadata), as well as the verification of retrieval results. A multimodal hybrid retrieval engine is constructed, achieving efficient collaborative retrieval of audio, text, and metadata through a three-level hierarchical index architecture. The audio modality uses the Wav2Clip model and PQ quantization technology to construct a compressed index, while the text modality uses the BGE-m3 model to achieve dense / sparse dual-path recall. Metadata is aligned with audio and text data through timestamps; specifically, cross-modal vector alignment technology is introduced to map audio and text to a unified semantic space. A dynamic routing engine is employed, automatically selecting the optimal retrieval path based on query features, providing an efficient infrastructure for multimodal content mining. High-precision retrieval such as "searching for text by audio" and "locating audio by text" can be achieved. This solves the core problems of traditional large models in audio content processing, such as high illusion rate, lack of domain specificity, and untraceable results.

[0174] This innovative approach aligns multi-path retrieval results and then uses a dynamic nonlinear fusion mechanism for fusion and reordering, achieving adaptive optimization and integration of cross-modal results. This solution achieves a breakthrough improvement in accuracy, credibility, and interpretability in podcast question-and-answer scenarios.

[0175] Furthermore, it achieves innovative visualization, intuitively displaying response information and associated timestamps for inquiries. It can directly provide feedback on the original slice content, as well as contextual content, summary content, citation content, and explanatory content, achieving multi-dimensional information feedback. Moreover, it can also provide comparative analysis, helping users quickly and comprehensively understand the perspectives of various parties in the relevant field.

[0176] Furthermore, a timeline is used to visually represent the overall structure, making the abstract temporal relationships between related content concrete. Jumps triggered by timestamps enable dynamic connections between content. In addition, a source tracing function is provided, making the origin of the response transparent. Users can directly view relevant evidence supporting their responses, providing credibility and helping them obtain effective information intuitively and accurately.

[0177] This solution, through rigorous generation control and innovative visual interaction, constructs a more professional and user-friendly search and question-and-answer process, enabling podcast content to evolve from traditional "passive listening" to "active exploration." Specifically, it transforms traditional one-way listening into an interactive, verifiable, and extensible knowledge exploration experience, setting a new standard for the intelligent application of audio knowledge bases.

[0178] This invention also provides a question-answering device, which can be used to implement any of the above-described method embodiments, and will not be repeated hereafter. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0179] This embodiment provides a question-and-answer device, such as... Figure 6 As shown, it includes: The response module is used to respond to received query requests; The generation module is used to generate response content corresponding to the query request; The response content includes response information used to reply to the query request, as well as a corresponding timestamp identifier, wherein the timestamp identifier is used to locate the slice information associated with the response information in time sequence.

[0180] The question-and-answer device provided in the embodiments of the present invention can execute the question-and-answer method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0181] The question-answering device in this embodiment generates a response to a received query request. This response includes reply information and a timestamp, where the timestamp is used to locate the slice information associated with the reply information in time sequence. This solution can respond to query requests and generate a timestamp containing reply information and the associated slice information, satisfying the user's need for proactive exploration and providing targeted, multi-dimensional response content.

[0182] Further functional descriptions of the above modules are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0183] Please see Figure 7 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. The electronic device can be a client or a server. The electronic device includes a memory, a processor, and a communication module. The memory, processor, and communication module are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.

[0184] The memory is used to store computer programs or data. Memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc.

[0185] The processor is used to read / write data or programs stored in the memory and to execute the question-and-answer method provided in any embodiment of the present invention.

[0186] The communication module is used to establish communication connections between electronic devices and other communication terminals via a network, and to send and receive data via the network.

[0187] It should be understood that, Figure 7 The structure shown is only a schematic diagram of an electronic device; the electronic device may also include components that are larger than those shown. Figure 7 The more or fewer components shown, or having the same Figure 7 The different configurations shown.

[0188] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from memory. When the computer program is executed by a processor, it performs the functions defined in the question-and-answer method of the embodiments of the present invention.

[0189] Figure 7The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0190] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory.

[0191] It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implement the question-and-answer method shown in the above embodiments.

[0192] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0193] In the embodiments provided by this invention, it should be understood that the disclosed apparatus and method can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0194] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0195] Furthermore, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0196] It should be noted that if the functionality is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0197] The above are merely embodiments of the present invention and are not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A question-and-answer method, characterized in that, The method includes: In response to a received query request, generate a response content corresponding to the query request; The response content includes response information for replying to the query request, and a corresponding timestamp identifier, wherein the timestamp identifier is used to locate the slice information associated with the response information in time sequence.

2. The question-and-answer method according to claim 1, characterized in that, The slice information includes the original slice content corresponding to the query request, and the method further includes: In response to a first trigger command for the timestamp identifier, the original slice content is fed back.

3. The question-and-answer method according to claim 2, characterized in that, The slice information also includes contextual content, summary content, citation content, and / or explanatory content corresponding to the original slice content, and the method further includes: In response to a second trigger command for the timestamp identifier, the context content, summary content, citation content, and / or explanatory content are fed back.

4. The question-and-answer method according to claim 2, characterized in that, The response content also includes a comparative analysis of the query request and the original slice content, and the method further includes: In response to a third trigger command for the timestamp identifier, the comparative analysis content is fed back.

5. The question-and-answer method according to any one of claims 2-4, characterized in that, The response content also includes a timeline and derived slice content arranged sequentially along the timeline. The derived slice content is associated with the original slice content. The timestamp identifier is configured as a movable control located on the timeline. The method further includes: In response to the timestamp being moved to the position corresponding to the derived slice content, the derived slice content is fed back.

6. The question-and-answer method according to claim 5, characterized in that, After the step of responding to the timestamp being moved to the position corresponding to the derived slice content and feeding back the derived slice content, the method further includes: In response to a fourth trigger command for the timestamp identifier, evidence information regarding the content of the derived slice is fed back.

7. The question-and-answer method according to claim 2, characterized in that, The step of generating a response to a received query request includes: The query request is used to search the corresponding database to determine the corresponding target search results, wherein the target search results correspond to the original slice content; The response content is generated based on the target search results.

8. The question-and-answer method according to claim 7, characterized in that, The steps of retrieving the corresponding target search results from the relevant database based on the query request include: The query request is subjected to intent analysis to determine the corresponding query intent, and the corresponding retrieval strategy is determined based on the query intent; The database is searched according to the determined search strategy to determine the target search results.

9. The question-and-answer method according to claim 8, characterized in that, The response content is generated by a generative large language model. The steps for generating the response content based on the target retrieval results include: The target search results are injected with corresponding prompt words using search enhancement generation technology; The prompt word instruction is input into the generative large language model to generate the response content.

10. The question-and-answer method according to claim 8, characterized in that, The steps of performing intent analysis on the query request to determine the corresponding query intent, and determining the corresponding retrieval strategy based on the query intent, include: The query request is input into the corresponding intent recognition model to determine the intent type of the query intent, and the corresponding intent parameters are extracted according to the intent type; The retrieval strategy is determined based on the intent type and the intent parameters. The retrieval strategy includes a retrieval mode that matches the query intent and a result reordering weight corresponding to the retrieval mode. The retrieval mode includes a text retrieval mode corresponding to factual query requests, a vector retrieval mode corresponding to opinion-based query requests, and / or a graph relationship retrieval mode corresponding to resource request query requests.

11. The question-and-answer method according to claim 10, characterized in that, The method is applied to digital media products containing audio, and the query request includes text-based query requests and / or audio-based query requests for the corresponding digital media content. The database includes: a vector database mapped to the same semantic space based on text data, audio data, and metadata; a text database based on text data and metadata; and / or an audio database based on audio data. The text data, audio data, and metadata are all configured with corresponding timestamps and correspond to the same digital media content based on the timestamps. The slice information is obtained by selecting segments of the digital media content based on the timestamp identifier. The metadata includes at least one of the following: the name, publisher, publication time, category tags, audio stream, notes, description, comments, citations, and resource links of the digital media content.

12. The question-and-answer method according to claim 11, characterized in that, The steps for determining the target search results by performing a search in the database according to the determined search strategy include: In response to the text-based query request, the text-based query request is converted into sparse vectors and dense vectors, and vector retrieval is performed in the vector database based on the sparse vectors and the dense vectors to determine the corresponding text vector retrieval results; and, Based on the text query request, a text search is performed in the text database to determine the corresponding text search results.

13. The question-and-answer method according to claim 12, characterized in that, The steps for determining the target search results by performing a search in the database according to the determined search strategy include: In response to the audio query request, the audio query request is converted into an audio vector, and a vector retrieval is performed in the vector database based on the audio vector to determine the corresponding audio vector retrieval result; The audio query request is subjected to speech recognition to convert it into a corresponding text retrieval request, and a text retrieval is performed in the text database based on the text retrieval request to determine the corresponding translated retrieval results; and / or, Based on the audio query request, an audio search is performed in the audio database to determine the corresponding audio search results.

14. The question-and-answer method according to claim 13, characterized in that, The step of performing a search in the database according to the determined search strategy to determine the target search results further includes: The text vector retrieval results, the text retrieval results, the audio vector retrieval results, the translated retrieval results, and / or the audio retrieval results are fused to obtain corresponding fused retrieval results; The target search result is determined in the fused search results based on the reordering weights of the results.

15. The question-and-answer method according to claim 14, characterized in that, The steps for fusing the text vector retrieval results, the audio vector retrieval results, the translated retrieval results, and / or the audio retrieval results to obtain corresponding fused retrieval results include: Based on the timestamps of the corresponding text data, audio data, and the metadata, the text vector retrieval results, the text retrieval results, the audio vector retrieval results, the translated retrieval results, and / or the audio retrieval results are timestamp aligned and deduplicated to obtain the corresponding fused retrieval results.

16. The question-and-answer method according to claim 14, characterized in that, The steps for determining the target search result in the fused search results based on the reordering weights of the results include: A non-linear weighting mechanism is adopted to reorder the fused search results based on the result reordering weights to obtain the corresponding reordered search results.

17. The question-and-answer method according to claim 16, characterized in that, The step of reordering the fused search results based on the result reordering weights using a non-linear weighting mechanism to obtain the corresponding reordered search results further includes: The slice content corresponding to the reordered search results is subjected to related content expansion processing, and then subjected to secondary reordering and secondary deduplication processing to obtain the target search results.

18. The question-and-answer method according to claim 9, characterized in that, The prompt word instructions include: The query request and the target search results; Instructions for setting roles in the generative large language model; Instructions for selecting the target search results; Output requirements instructions for the content of the response; Inference restriction instructions for the content of the response.

19. The question-and-answer method according to claim 18, characterized in that, The step of generating the response content based on the target retrieval results further includes: Perform content verification and / or source tracing verification on the generated response content, and output the response content after the verification is successful.

20. The question-and-answer method according to claim 11, characterized in that, The text data in the database includes structured fields, unstructured data, and reference relationships corresponding to the digital media content. After being vectorized by a text encoder, the text data is stored in the vector database in the form of sparse text vectors and dense text vectors.

21. The question-and-answer method according to claim 11, characterized in that, After optimization, the audio data is encoded into relatively aligned text vectors and audio vectors by a text-audio dual encoder. The text vectors are stored in the text database, and the audio vectors are stored in the vector database. The optimization process includes: timbre adaptation optimization, long-term context modeling, and / or scene-aware feature enhancement.

22. The question-and-answer method according to claim 21, characterized in that, The audio data originates from the pre-processed original audio, and the pre-processing includes: noise reduction processing, silent segment removal processing, speech enhancement processing, and / or channel conversion processing.

23. A question-and-answer device, characterized in that, The device includes: The response module is used to respond to received query requests; A generation module is used to generate response content corresponding to the query request; The response content includes response information for replying to the query request, and a corresponding timestamp identifier, wherein the timestamp identifier is used to locate the slice information associated with the response information in time sequence.

24. An electronic device, characterized in that, It includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the question-answering method according to any one of claims 1 to 22.

25. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the question-and-answer method according to any one of claims 1 to 22.

26. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the question-and-answer method according to any one of claims 1 to 22.