Radio content generation method and device based on space-time scene association and agent memory, equipment and medium

CN122388271BActive Publication Date: 2026-08-21XIANGJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610859710.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-08-21
Estimated Expiration
2046-06-15

AI Technical Summary

Technical Problem

[0004]然而,现有的基于传统用户画像的方案难以捕获用户在动态时空场景下的偏好迁移,这导致推荐逻辑在跨场景间缺乏连续性与自适应能力

Benefits of technology

首先,通过根据外部时空特征和用户交互数据更新多层智能体记忆库,动态收集并维护分层的记忆状态数据,实现了对环境状态变化与用户历史行为特征的结构化沉淀;其次,通过大语言模型对多层智能体记忆库进行意图解析,得到用户即时意图、记忆推荐权重以及语义探索权重,该步骤量化了历史偏好延续与新意图探索的分配比例,为后续资源召回提供了明确的定量依据;接着,根据记忆推荐权重进行记忆定向召回,并根据语义探索权重进行语义探索召回,得到带有溯源标签的内容候选池,保障了初步召回的音频兼具熟悉度与多样性,且推荐来源可清晰回溯;随后,基于从多层智能体记忆库中提取的全局负向标签和场景负向标签对内容候选池进行双重过滤,得到待排序候选列表,精准隔离了全局排斥以及当前具体场景下不适宜的不良体验内容;在此基础上,基于记忆推荐权重、语义探索权重以及用户即时意图对待排序候选列表中的候选内容进行多维加权打分,得到推荐内容序列,该过程统筹兼顾了多维度的得分指标进行综合评价与排序,实现了推荐资源在历史偏好延续与新颖内容探索间的平滑过渡与优选;最后,根据溯源标签、记忆推荐权重以及语义探索权重生成交互导语,并将交互导语与推荐内容序列对应的音频进行组合,得到最终的电台内容,基于内容召回来源及推荐逻辑动态生成对应的语音并进行拼接播报,保障了文本事实与推荐逻辑的一致性;综合上述多维度的感知、分析与匹配流程,本申请能够提升智能电台的个性化陪伴感,为用户提供跨场景连贯且情感自然契合的沉浸式收听体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122388271B_ABST
    Figure CN122388271B_ABST
Patent Text Reader

Abstract

The application discloses a radio content generation method and device based on space-time scene association and agent memory, equipment and medium, relates to the technical field of media content recommendation, and the method comprises the steps that external space-time characteristics and interactive data are updated to obtain a multi-layer agent memory library composed of work, scene and semantic layer; a large model is called to analyze the context, output user immediate intention, memory recommendation and semantic exploration weight; according to the weight proportion, the directional and exploration double-way recall are executed, and the candidate pool with the traceability label is constructed; the global and scene negative label of the memory library are extracted to filter the content, and the to-be-sequenced list is generated; multi-dimensional weighted sequencing is performed combined with double-way weight and intention to determine the recommendation sequence; the interactive lead-in is generated by matching the generation strategy, and after voice conversion, the radio content is obtained by merging with the audio. The application can improve the personalized accompaniment of the intelligent radio, and provide the user with immersive listening experience that is coherent across scenes and naturally fits the emotions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of media content recommendation technology, and in particular to a method, apparatus, device and medium for generating radio content based on spatiotemporal scene association and intelligent agent memory. Background Technology

[0002] With the widespread adoption of smart devices and IoT technologies, smart radio, as a companion audio service platform, is gradually transforming from simple content distribution to personalized and intelligent audio companion services. Users' listening intentions vary significantly across different scenarios (such as commuting, work, and bedtime), which places higher demands on radio systems in terms of multimodal context awareness and long-term user profile construction.

[0003] Existing intelligent radio audio recommendation systems primarily rely on traditional user profiling techniques when processing user requests. These systems typically collect data such as users' historical listening records, likes, and favorites to build static user preference models. When recommending content, the system matches audio files from its media content library based on these static features. Simultaneously, some systems introduce basic negative feedback mechanisms to filter content that users explicitly indicate they dislike. At the interaction level, some systems generate an interactive introduction using text-to-speech (TTS) technology, based on the recommendation results and a pre-set simple template, before audio playback.

[0004] However, existing solutions based on traditional user profiles struggle to capture the shifts in user preferences across dynamic spatiotemporal scenarios, resulting in a lack of continuity and adaptability in recommendation logic across different scenarios. Secondly, the lack of a smooth transition between continuing historical preferences and exploring new intentions significantly reduces the user's immersive experience. Furthermore, existing negative feedback mechanisms often lack scenario isolation capabilities and have low fault tolerance. Finally, in generating interactive introductions, the lack of a natural interactive loop often leads to a factual mismatch between the generated text and the actual playback content, resulting in a lack of interactive consistency, which severely weakens the system's sense of companionship and credibility. Therefore, how to enhance the personalized companionship of smart radio and provide users with a coherent and emotionally resonant immersive listening experience across scenarios has become an urgent problem to be solved. Summary of the Invention

[0005] The purpose of this application is to provide a method, device, equipment, and medium for generating radio content based on spatiotemporal scene association and intelligent agent memory, aiming to solve the technical problem of how to enhance the personalized companionship of intelligent radio and provide users with an immersive listening experience that is coherent across scenes and emotionally natural.

[0006] To achieve the above objectives, this application proposes a radio content generation method based on spatiotemporal scene association and agent memory, the method comprising: The multi-layered intelligent agent memory bank is updated based on external spatiotemporal characteristics and user interaction data. The intent is parsed by the multi-layered agent memory database using a large language model to obtain the user's immediate intent, memory recommendation weight, and semantic exploration weight. Based on the memory recommendation weight, memory-oriented recall is performed, and based on the semantic exploration weight, semantic exploration recall is performed to obtain a content candidate pool with source tracing tags; The content candidate pool is filtered based on the global negative labels and scene negative labels extracted from the multi-layer intelligent agent memory library to obtain a candidate list to be sorted. Based on the memory recommendation weight, the semantic exploration weight, and the user's immediate intent, the candidate content in the unsorted candidate list is scored using a multi-dimensional weighted score to obtain a recommended content sequence. An interactive introduction is generated based on the source tag, the memory recommendation weight, and the semantic exploration weight. The interactive introduction is then combined with the audio corresponding to the recommended content sequence to obtain radio content.

[0007] In one embodiment, the multi-layered intelligent agent memory bank includes a working memory layer, a contextual memory layer, and a semantic memory layer; The step of updating the multi-layer intelligent agent memory bank based on external spatiotemporal characteristics and user interaction data includes: The external spatiotemporal features are discretized to obtain the current scene data, and the original dialogue text sequence and listening behavior feedback are parsed from the user interaction data. The current scene data and the original dialogue text sequence are stored in the working memory layer; Update the scene score of the corresponding scene in the scene memory layer based on the listening behavior feedback; When the scene integral is lower than a preset negative threshold, a scene negative label and a global negative label are generated; The original dialogue text sequence is semantically compressed using a large language model to generate a dialogue summary, and features are extracted from the dialogue summary using the large language model to obtain the user's long-term interests. The scene negative label, the dialogue summary, and the scene score are stored in the context memory layer, and the global negative label and the user's long-term interests are stored in the semantic memory layer, thus completing the update.

[0008] In one embodiment, the step of parsing the multi-layer agent memory database using a large language model to obtain the user's immediate intent, memory recommendation weight, and semantic exploration weight includes: Retrieve related scenes from the scene memory layer that have a feature overlap of greater than or equal to a preset overlap threshold with the current scene data in a preset dimension; Extract the dialogue summary corresponding to the associated scenario, and concatenate the dialogue summary into historical context according to the time sequence; The preset weighted inference rules, the historical context, the current scene data, and the original dialogue text sequence are used to construct structured prompt words; The structured prompts are input into the large language model for quantitative reasoning calculations to obtain the user's immediate intent, memory recommendation weight, and semantic exploration weight.

[0009] In one embodiment, the step of performing memory-oriented recall based on the memory recommendation weight and semantic exploration recall based on the semantic exploration weight to obtain a content candidate pool with source tracing tags includes: The first recall quantity corresponding to the memory path is determined based on the memory recommendation weight, and the first candidate content with a scene score greater than or equal to a preset score threshold is retrieved from the radio content library based on the first recall quantity. Attach memory tags to the first candidate content to obtain memory-oriented recall results; The second recall quantity corresponding to the exploration path is determined based on the semantic exploration weight, and the user's immediate intent is converted into an intent vector; In the radio content library, based on the second recall quantity, a second candidate content matching the intent vector is recalled using an approximate nearest neighbor retrieval algorithm; Add semantic tags to the second candidate content to obtain semantic exploration recall results; The memory-oriented recall results and the semantic exploration recall results are deduplicated and merged to obtain a content candidate pool with source tags.

[0010] In one embodiment, the multi-layered intelligent agent memory bank includes a working memory layer, a contextual memory layer, and a semantic memory layer; The step of filtering the content candidate pool based on global negative labels and scene negative labels extracted from the multi-layer agent memory to obtain a candidate list to be sorted includes: Extract global negative labels from the semantic memory layer; Based on the global negative tags, candidate content in the content candidate pool is matched and eliminated to obtain a preliminary filtered content pool; Extract current scene data from the working memory layer, and extract scene negative labels corresponding to the current scene data from the context memory layer; Based on the negative tags of the scenario, candidate content in the initial filtered content pool is matched and eliminated to obtain a candidate list to be sorted.

[0011] In one embodiment, the step of performing multi-dimensional weighted scoring on candidate content in the candidate list to be ranked based on the memory recommendation weight, the semantic exploration weight, and the user's immediate intent to obtain a recommended content sequence includes: When the candidate content in the unsorted candidate list carries a memory tag, the maximum scene score under the associated scene is extracted as the scene preference score, and when the memory tag is not carried, the scene preference score is set to a preset initial score. The user's long-term interests are extracted from the semantic memory layer of the multi-layer intelligent agent memory library, and the first similarity between the user's long-term interests and the content semantic features of the candidate content is calculated to obtain the interest preference score. Calculate the second similarity between the user's immediate intent and the semantic features of the content to obtain the intent exploration score; The memory path score is obtained by weighting and summing the scene preference score and the interest preference score based on the memory recommendation weight. Multiply the semantic exploration weight by the intent exploration score to obtain the exploration path score, and sum the memory path score with the exploration path score to obtain the multidimensional weighted total score; Based on the multidimensional weighted total score, the candidate content is sorted in descending order and its position is truncated to obtain the recommended content sequence.

[0012] In one embodiment, the step of generating an interactive introduction based on the source tag, the memory recommendation weight, and the semantic exploration weight, and combining the interactive introduction with the audio corresponding to the recommended content sequence to obtain radio content includes: Based on the source tag, the memory recommendation weight, and the semantic exploration weight, a corresponding lead generation strategy is matched in the preset strategy library; The sentiment type vector and sentiment intensity coefficient are determined based on the aforementioned lead generation strategy; Extract the attribute information of the recommended content sequence, and construct the prompt word template by combining the attribute information, the user's immediate intent, and the current scene data in the working memory layer of the multi-layer intelligent agent memory bank; The prompt word template is input into the large language model to generate text, resulting in interactive introductory text. Based on the emotion type vector and the emotion intensity coefficient, the interactive introductory text is converted into an introductory speech stream; The introductory audio stream is concatenated with the audio corresponding to the recommended content sequence and then output to obtain the radio content.

[0013] Furthermore, to achieve the above objectives, this application also proposes a radio content generation device based on spatiotemporal scene association and agent memory, the device comprising: The memory update module is used to update the multi-layer intelligent agent memory bank based on external spatiotemporal characteristics and user interaction data; The intent parsing module is used to parse the intent of the multi-layer intelligent agent memory database through a large language model to obtain the user's immediate intent, memory recommendation weight, and semantic exploration weight. The dual-path recall module is used to perform memory-oriented recall based on the memory recommendation weight and semantic exploration recall based on the semantic exploration weight, so as to obtain a content candidate pool with source tracing tags. The negative filtering module is used to filter the content candidate pool based on global negative labels and scene negative labels extracted from the multi-layer intelligent agent memory library to obtain a candidate list to be sorted. The multi-dimensional scoring module is used to perform multi-dimensional weighted scoring on the candidate content in the candidate list to be sorted based on the memory recommendation weight, the semantic exploration weight, and the user's real-time intent, so as to obtain a recommended content sequence. The content synthesis module is used to generate an interactive introduction based on the source tag, the memory recommendation weight, and the semantic exploration weight, and to combine the interactive introduction with the audio corresponding to the recommended content sequence to obtain radio content.

[0014] Furthermore, to achieve the above objectives, this application also proposes a radio content generation device based on spatiotemporal scene association and agent memory. The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the radio content generation method based on spatiotemporal scene association and agent memory as described above.

[0015] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the radio content generation method based on spatiotemporal scene association and agent memory as described above.

[0016] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the radio content generation method based on spatiotemporal scene association and agent memory as described above.

[0017] One or more technical solutions proposed in this application have at least the following technical effects: First, by updating the multi-layered agent memory bank based on external spatiotemporal features and user interaction data, hierarchical memory state data is dynamically collected and maintained, achieving structured accumulation of environmental state changes and user historical behavior characteristics. Second, intent parsing is performed on the multi-layered agent memory bank using a large language model to obtain the user's immediate intent, memory recommendation weight, and semantic exploration weight. This step quantifies the allocation ratio between historical preference continuation and new intent exploration, providing a clear quantitative basis for subsequent resource retrieval. Next, memory-oriented retrieval is performed based on the memory recommendation weight, and semantic exploration retrieval is performed based on the semantic exploration weight, resulting in a content candidate pool with source-tracing tags. This ensures that the initially retrieved audio has both familiarity and diversity, and the recommendation source can be clearly traced back. Subsequently, the content candidate pool is double-filtered based on global negative tags and scene negative tags extracted from the multi-layered agent memory bank to obtain a candidate list to be ranked, accurately isolating global exclusion. This includes identifying unsuitable or negative user experiences in specific scenarios. Based on this, a multi-dimensional weighted scoring system is applied to candidate content in the ranking list, incorporating memory recommendation weights, semantic exploration weights, and the user's immediate intent. This results in a recommended content sequence, comprehensively evaluating and ranking multiple scoring metrics to achieve a smooth transition and optimal selection of recommended resources between historical preference continuity and novel content exploration. Finally, interactive introductory text is generated based on source tags, memory recommendation weights, and semantic exploration weights. This introductory text is then combined with the audio corresponding to the recommended content sequence to obtain the final radio content. The corresponding audio is dynamically generated and spliced ​​together based on the content recall source and recommendation logic, ensuring consistency between textual facts and recommendation logic. By integrating these multi-dimensional perception, analysis, and matching processes, this application enhances the personalized companionship of smart radio, providing users with a cross-scenario, coherent, and emotionally resonant immersive listening experience. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating an embodiment of the radio content generation method based on spatiotemporal scene association and agent memory in this application. Figure 2This is a flowchart illustrating Embodiment 2 of the radio content generation method based on spatiotemporal scene association and agent memory in this application. Figure 3 A simplified flowchart illustrating the radio content generation method based on spatiotemporal scene association and agent memory provided in Embodiment 2 of this application; Figure 4 This is a schematic diagram of the module structure of the radio content generation device based on spatiotemporal scene association and intelligent agent memory in an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the radio content generation method based on spatiotemporal scene association and agent memory in the embodiments of this application.

[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0022] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0024] It should be noted that the executing entity of this application embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or radio content generation system capable of realizing the above functions. The following description uses a radio content generation system as an example to illustrate this embodiment and the subsequent embodiments.

[0025] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0026] Based on this, embodiments of this application provide a method for generating radio content based on spatiotemporal scene association and agent memory, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the radio content generation method based on spatiotemporal scene association and agent memory of this application.

[0027] In this embodiment, the radio content generation method based on spatiotemporal scene association and agent memory includes steps S10 to S60: Step S10: Update the multi-layer intelligent agent memory bank based on external spatiotemporal characteristics and user interaction data; It should be noted that external spatiotemporal features can include the device's GPS coordinates, current timestamp, and weather conditions obtained from a meteorological interface. User interaction data refers to the behavioral records generated by users when using the radio application, including completion of playback, skipping, playback duration, and dialogue commands input via text or voice. A multi-layered agent memory refers to a hierarchical storage architecture consisting of a working memory layer, a contextual memory layer, and a semantic memory layer, used for short-term state recording, historical scene snapshot storage, and long-term interest extraction, respectively. The working memory layer is the memory level used to store real-time multimodal state data of the current session. Its stored content can include real-time GPS coordinates, region, weather, current time, time period, periodic time (weekdays / holidays), and the original text sequence of the current dialogue and the user's real-time intent. The contextual memory layer is the memory level used to store historical interaction content and summaries in specific scenarios. The content data stored in the user's specific scenario can include content ID, interaction time, user interaction feedback, listening duration, and scenario points. The semantic memory layer refers to the memory level used to store global preferences summarized from long-term interactions. Its contents may include global preference scores, user long-term interests, content IDs of scene negative labels, and content IDs of global negative labels.

[0028] Understandably, the process involves acquiring the latitude and longitude data currently fed back by the mobile device, converting it into a business district or administrative division area through reverse geocoding; extracting the system time and mapping it to the corresponding time period (e.g., weekday morning); and simultaneously acquiring the current weather conditions (e.g., sunny). This information, along with the user's current dialogue command text, is directly stored in the working memory layer. Subsequently, increments are calculated based on the user's playback behavior statistics. For example, completing a single audio track increases the corresponding scene score by 1; fast-forwarding or skipping tracks decreases the scene score by 1. Records containing the scores and dialogue text are stored in the contextual memory layer. Further, the dialogue command text is input into a large language model to generate a dialogue summary, and the long-term preference features in the semantic memory layer are updated based on the historical accumulated score values. This step achieves hierarchical structured storage of multi-dimensional contextual data, enabling the system to capture and quantify environmental changes and user behavior in dynamic scenarios, providing data support for subsequent cross-scene preference perception.

[0029] Step S20: The intent is parsed by the multi-layer intelligent agent memory database through the large language model to obtain the user's immediate intent, memory recommendation weight and semantic exploration weight; It should be noted that the user's immediate intent refers to the specific request for a particular style or type of audio content expressed by the user in the current conversation. The memory recommendation weight refers to the proportion of previously preferred content assigned to the user when allocating recall resources, ranging from 0.0 to 1.0. The semantic exploration weight can be the proportion allocated to entirely new or previously unexplored content, ranging from 0.0 to 1.0, and its sum with the memory recommendation weight is strictly equal to 1.0.

[0030] Understandably, the historical records are traversed in the contextual memory layer, comparing time periods, weather, and location information. If the number of feature dimensions matching the historical record and the current state is greater than or equal to a preset overlap threshold (set to 2 to ensure sufficient relevance of the scene), a dialogue summary for the corresponding historical scene is extracted. The extracted historical dialogue summary, current environmental state information, and user commands are concatenated into a prompt text, which is then input into the large language model. The large language model outputs JSON-formatted text based on thought chain reasoning. For example, when a user inputs the command "want to hear something I haven't heard before," the large language model outputs text containing target intent words and outputs a memory recommendation weight of 0.2 and a semantic exploration weight of 0.8. This step utilizes the natural language understanding capabilities of the large language model to transform ambiguous contextual information into definite intent labels and quantified recommendation allocation ratios, solving the resource allocation problem between the continuation of historical preferences and the exploration of novel content.

[0031] Step S30: Perform memory-oriented recall based on the memory recommendation weight, and perform semantic exploration recall based on the semantic exploration weight to obtain a content candidate pool with source tracing tags; It should be noted that memory-oriented recall refers to the process of retrieving highly-rated audio from historically related scene data as recommendation candidates. Semantic exploration recall can be a process of retrieving novel audio similar to the current intent from a global media library using vector space matching technology. Source tags are attribute identifiers used to identify the content recall path, including both memory tags and semantic tags. The content candidate pool refers to a set of radio station audios to be ranked after initial screening.

[0032] Understandably, the system's base recall count is set at 100. Based on the weight values ​​obtained in step S20, the audio is allocated as follows: if the memory recommendation weight is 0.6, then from the similar scenes matched in the context memory layer, the audio is sorted in descending order of context score, and the top 60 audio clips are extracted and labeled with memory tags. Similarly, the semantic exploration weight is 0.4. The user's immediate intent text is converted into a feature vector using the Word2Vec model. In the total radio content library, the Approximate Nearest Neighbor (ANN) algorithm is used to find the top 40 audio clips with the highest cosine similarity, and semantic tags are added to these audio clips. The two sets of audio clips are merged, and duplicates are removed to form a content candidate pool. This step integrates relevance recommendation based on historical memory and novelty recommendation based on semantic generalization discovery, ensuring that the recalled content has both familiarity and diversity, and recording the source logic of each audio clip through tags.

[0033] Step S40: Filter the content candidate pool based on the global negative labels and scene negative labels extracted from the multi-layer intelligent agent memory library to obtain a candidate list to be sorted. It should be noted that global negative tags refer to content that users consistently dislike over long-term cross-scenario interactions, resulting in its overall blocking by the system. Scenario negative tags can be content that users dislike only in a specific environment or time period. The candidate list to be sorted refers to the audio collection retained after initial recall and removal of content that provides a negative experience.

[0034] Understandably, the system backend sets two evaluation thresholds: a global negative label is generated when the audio's global preference score falls below a negative threshold (set to -5 points); a scene negative label is generated when the audio's score in a specific scene falls below a single scene negative threshold (set to -3 points). The global negative label list is retrieved from the semantic memory layer, and the audio IDs in the content candidate pool are iterated through. If an audio carries a global negative label, it is directly removed from the pool. Subsequently, the current environment state is extracted from the working memory layer, and the corresponding scene negative label list is retrieved from the context memory layer. This is compared with the remaining audio, and audio with matching environment labels is removed again. Audio that has undergone two rounds of removal is retained for the next stage. This step establishes a dual isolation filtering mechanism, which not only avoids content that users generally dislike but also finely filters audio that is unsuitable for playback in specific environments, reducing the recommendation error tolerance.

[0035] Step S50: Based on the memory recommendation weight, the semantic exploration weight, and the user's immediate intent, perform multi-dimensional weighted scoring on the candidate content in the candidate list to be sorted to obtain a recommended content sequence; It should be noted that multidimensional weighted scoring refers to the process of comprehensively considering multiple environmental and content relevance indicators to calculate a final comprehensive score for each element in the set. The recommended content sequence can be a list of audio files to be played, arranged from highest to lowest based on the final score.

[0036] Understandably, for audio tracks with memory tags in the list, their historical scene integral is extracted and multiplied by the corresponding time decay factor (this factor is calculated based on the time difference between the current time and the last playback time; the larger the time difference, the greater the decay value, with a value range of 0 to 1). The result is added to the cosine similarity between the audio track and the user's long-term interests, and then multiplied by the memory recommendation weight to obtain the memory score. For audio tracks with semantic tags, the dot product similarity between the user's immediate intent vector and the audio track's feature vector is calculated and multiplied by the semantic exploration weight to obtain the exploration score. The memory score and exploration score of each audio track are summed to obtain a comprehensive score for each audio track. These tracks are then sorted in descending order of their comprehensive scores, and the top N audio tracks constitute the final playlist. This step integrates the user's long-term preferences, the timeliness of scene feedback, and current immediate needs, achieving a unified measurement and ranking of content from different recall sources through quantitative calculation.

[0037] Step S60: Generate an interactive guide based on the source tag, the memory recommendation weight, and the semantic exploration weight, and combine the interactive guide with the audio corresponding to the recommended content sequence to obtain radio content.

[0038] It should be noted that interactive introductions refer to the text and synthesized speech content that the system broadcasts to the user before playing the main audio, providing explanations and conveying a specific emotional tone. The radio content can be a data stream that ultimately appears to the user, containing a mix of audio and music clips.

[0039] Understandably, the system's built-in generation strategy is matched using conditional judgment logic. When the audio's source tag is a memory tag and the memory score is greater than the exploration score, the contextual resonance strategy is selected. Under this strategy, the corresponding audio's stylistic features and current weather information are extracted and input into the large language model, with text tone constraints set. The large language model outputs text similar to "On a rainy day, listen to some soothing light music I used to listen to." Subsequently, using an end-to-end text-to-speech (TTS) model with acoustic parameter control capabilities, combined with the emotional tension parameter set in the strategy (setting the intensity coefficient to 1.0 to represent a warm tone), the text is synthesized into an audio stream. Finally, this synthesized audio stream is appended to the beginning of the recommended content sequence for continuous playback. This step introduces a natural speech transition that matches the current recommendation logic before the audio officially plays, mitigating the abruptness of content switching and enhancing emotional transmission during human-computer interaction.

[0040] This embodiment establishes a closed-loop audio distribution system with spatiotemporal awareness by combining a hierarchical memory architecture with large language model reasoning. This solution decouples users' long-term interests from their immediate short-term needs, effectively balancing the continuation of historical habits with the exploration of unknown domains during the resource retrieval phase. Simultaneously, the dual negative feedback filtering mechanism and multi-strategy-driven voice prompt generation complement each other, improving the relevance of recommended content in different contexts, ensuring the logical coherence and emotional consistency of human-computer interaction, and providing users with intelligent audiovisual services that offer strong companionship and an immersive experience.

[0041] As an example, the multi-layered agent memory includes a working memory layer, a contextual memory layer, and a semantic memory layer. The step of updating the multi-layered agent memory based on external spatiotemporal features and user interaction data includes: discretizing the external spatiotemporal features to obtain current scene data, and parsing the original dialogue text sequence and listening behavior feedback from the user interaction data; storing the current scene data and the original dialogue text sequence in the working memory layer; updating the scene integral of the corresponding scene in the contextual memory layer based on the listening behavior feedback; generating scene negative labels and global negative labels when the scene integral is lower than a preset negative threshold; semantically compressing the original dialogue text sequence using a large language model to generate a dialogue summary, and extracting features from the dialogue summary using the large language model to obtain the user's long-term interests; storing the scene negative labels, the dialogue summary, and the scene integral in the contextual memory layer, and storing the global negative labels and the user's long-term interests in the semantic memory layer to complete the update.

[0042] It should be noted that discretization refers to the data processing operation of converting continuous numerical features into discrete category labels. Current scene data refers to the set of contextual features that possess environmental semantic information after discretization mapping. The original dialogue text sequence can be text entered by the user during the interactive session or an unprocessed word sequence converted from speech. Listening behavior feedback can be records of user interactions such as completing playback, skipping, and adding to favorites during audio playback. Scene score refers to a numerical indicator used to quantitatively evaluate the user's preference for audio content within a specific spatiotemporal environment. The preset negative threshold is the judgment boundary used to trigger the generation of negative feedback labels; its value is usually a negative integer (e.g., -3 or -5), and this value is set based on the user's historical bounce rate and tolerance statistics. Dialogue summary refers to a short text generated after extracting the core semantics of longer interactive text using a large language model. User long-term interests refer to stable preference features such as audio genres and topic classifications extracted from long-term historical interactions across scenarios.

[0043] In addition, the memory update mechanism occurs at the end of each round of dialogue, converting the user's interaction behavior in that round into a quantified increment (e.g., like +2, skip -1, complete +1, dislike -1) to update the corresponding scene score. The system uses a large language model to semantically compress the dialogue text temporarily stored in the working memory layer, generating a dialogue semantic summary, which is then stored synchronously with the interaction content in the context memory layer. The large model is then used to summarize the dialogue summaries in the context memory layer to form the user's long-term interests. Accumulated points... When the value drops to a preset first negative threshold, a negative label is assigned to the scene; when the content's global preference score across multiple scenes... When the value drops to a preset second negative threshold, a global negative label is assigned. At the same time, a time decay mechanism is introduced to increase the scene score of suspended content that has not been interacted with for a long time (also known as scene cooldown revival) to adapt to dynamic changes in preferences.

[0044] Understandably, the first step is to discretize the external spatiotemporal features. This involves acquiring the external spatiotemporal features reported by the device's underlying layer, mapping consecutive timestamps to corresponding time periods (e.g., weekday mornings, weekend nights), using a reverse geocoding interface to convert latitude and longitude values ​​into regional affiliation labels (e.g., commercial areas, residential areas), and mapping meteorological data to conventional weather categories (e.g., sunny days, rainy days). The current scene data is then obtained by combining these discrete categories. Simultaneously, user interaction data is parsed, extracting the input text or speech-transcribed text to form the original dialogue text sequence, and extracting the operation logs from the playback interface as listening behavior feedback. Next, the current scene data and the original dialogue text sequence are temporarily stored in the working memory layer, providing an immediate contextual environment for the current interaction session.

[0045] Then, user preferences are quantified to update the scene score. Corresponding quantification increment mapping rules are configured for listening behavior feedback; for example, completing a playback is mapped to adding 1 point, adding to favorites to add 2 points, and skipping playback is mapped to subtracting 1 point. The quantification increment is calculated based on the extracted specific listening behavior feedback, and this increment is added to the historical scene score matching the current scene in the scene memory layer to obtain the updated scene score.

[0046] Next, the negative isolation judgment logic is executed. The updated scene score is extracted, and it is determined whether the scene score is lower than the preset negative threshold. When the scene score drops below the preset negative threshold (e.g., -3 points), a scene negative label is generated for the corresponding content; if the content continues to receive negative evaluations in multiple rounds of interaction and drops to an even lower set threshold (e.g., -5 points), a global negative label is generated simultaneously.

[0047] Furthermore, abstract semantic information is extracted. The original dialogue text sequence in the working memory layer is input into a large language model. The natural language understanding capabilities of the large language model are used to remove redundant interjections and compress the contextual structure, generating a refined dialogue summary. Subsequently, the dialogue summary is input into the large language model again for deep feature induction, from which preference elements such as audio style and content theme are extracted to obtain the user's long-term interests.

[0048] Finally, the state synchronization of the memory bank is performed. The scene negative labels, dialogue summaries, and updated scene scores are stored in the corresponding storage nodes of the context memory layer; at the same time, the global negative labels and user long-term interests are stored in the semantic memory layer, completing the data update and flow of the multi-layer agent memory bank.

[0049] This example, through the steps outlined above, transforms unstructured underlying physical signals and natural language into hierarchical memory state features. Utilizing discretization mapping and quantized integral calculation, it achieves an objective depiction of dynamic spatiotemporal scenes and fine-grained interactive behaviors. Combined with the summarization and feature extraction capabilities of a large language model, it efficiently decouples immediate states, contextual snapshots, and stable preferences. The established dual negative judgment mechanism effectively isolates undesirable interactive experiences. The overall process ensures the continuity of data recording across scenarios, providing a complete data foundation for intent perception and personalized content retrieval.

[0050] As an example, the step of parsing the intent of the multi-layered agent memory database using a large language model to obtain the user's immediate intent, memory recommendation weight, and semantic exploration weight includes: retrieving related scenes in the context memory layer whose feature overlap with the current scene data in a preset dimension is greater than or equal to a preset overlap threshold; extracting the dialogue summary corresponding to the related scene, and concatenating the dialogue summary into a historical context according to a time sequence; constructing structured prompt words using preset weight inference rules, the historical context, the current scene data, and the original dialogue text sequence; and inputting the structured prompt words into the large language model for quantitative inference calculation to obtain the user's immediate intent, memory recommendation weight, and semantic exploration weight.

[0051] It should be noted that in this example, the preset dimension refers to the feature consideration direction used for scene similarity matching, mainly including time, location, and weather. Feature overlap refers to the number of two scene data points that have completely identical values ​​on the corresponding preset dimension. Related scenes refer to historical context records that are highly similar to the current environment in terms of feature representation. Historical context refers to long text information composed of multiple interaction records from related scenes strung together in chronological order. Preset weight inference rules refer to the logical boundary conditions that constrain the output proportion of the large language model, such as requiring the sum of the two output weight values ​​to be strictly 1.0 and accurate to one decimal place. Structured prompts refer to text instructions organized according to a specific format (such as JSON), containing logical rules, input data, and output format requirements. User immediate intent refers to the user's direct request for specific audio content expressed in the current session round. Memory recommendation weight refers to the quantified proportion assigned to historically preferred content in the subsequent recall stage. Semantic exploration weight refers to the quantified proportion assigned to globally unknown or novel content in the subsequent recall stage.

[0052] Understandably, the first step is to perform feature comparison in the context memory layer. This involves iterating through the historical scene records stored in the context memory layer, extracting the time, location, and weather dimensions from the current scene data, and performing attribute matching with the corresponding dimensions of the historical scene records. The number of dimensions where both sets of data have identical values ​​on these preset dimensions is counted, yielding the corresponding feature overlap. Historical scene records with feature overlap values ​​greater than or equal to a preset overlap threshold are then selected and extracted as associated scenes.

[0053] Next, the historical context information is organized. The stored dialogue summaries corresponding to each selected related scenario are obtained. Based on the chronological order of the related scenarios, the extracted dialogue summaries are concatenated into a continuous historical context.

[0054] Next, assemble the data template for model input. Obtain the preset weighted inference rules configured in the system backend. For example, the rules may explicitly require the output to include the inference process, intent description, and two numerical weights, and limit the sum of the two weights to 1.0. Then, fill in the fields according to the system's preset template skeleton using the obtained preset weighted inference rules, the concatenated historical context, the current scene data, and the original dialogue text sequence obtained from the current session, constructing structured prompt words.

[0055] Finally, the constructed structured prompts are input into the large language model for quantitative inference calculation. The large language model combines historical contextual information with the environmental state of the current scene data to analyze the deep semantics of the original dialogue text sequence, outputting a text result representing the user's current request as the user's immediate intent. Simultaneously, the large language model strictly adheres to the constraints of preset weighted inference rules, outputting two specific floating-point values ​​for the current context, serving as the memory recommendation weight and semantic exploration weight, respectively.

[0056] To eliminate the uncontrollability of the output of the large language model, historical contextual memory and real-time working memory are combined to construct structured prompt words, which are then input into the large language model for quantitative reasoning using the mind chain technique. The specific prompt word construction module includes: (1) Task and format constraints: LLM is defined as an intent parsing and weight allocation engine, and the output format is required to be standardized JSON, including the reasoning process, core intent, and memory recommendation weights ( ) and semantic exploration weights ( (1) **Constraints:** Clearly define that the two weight values ​​must be accurate to one decimal place and that their sum must be strictly 1.0. (2) **Input Context:** Access the previously assembled historical context summary, as well as the scene data and interaction context containing the user's current dialogue. (3) **Inference Rule Boundaries:** Set clear weight judgment logic. For example, if the user's immediate input contains a clear request for new information, then a strict constraint is imposed. .

[0057] Example of structured cue word construction: [Role and Task]: You are the intent parsing and traffic allocation engine for an intelligent radio station. Please analyze the user's real-time context and historical memory, and output the memory recommendation weight accurate to one decimal place. ) and semantic exploration weights ( And the sum of the two is strictly 1.0.

[0058] [Real-time Working Memory]: Scene Data: Current Time (Weekend 22:30), Weather (Rainy Day), Location (Yuelu District). Interaction Context: Input Command: "Play some relaxing music".

[0059] [Historical Context Memory]: Historical Interaction Content and Summary: Users have repeatedly played and saved light jazz and ambient white noise style tracks in related scenarios.

[0060] [Inference Rule Constraints]: If the user's immediate input contains an explicit request for novelty (such as "listen to something you haven't heard of before"), enforce the constraint. .

[0061] [Output Format Requirements]: Output must strictly adhere to JSON format and must not contain any irrelevant characters. It must include the following fields: "reasoning": a description of the reasoning process; "intent": the user intent behind the reasoning; ": Memory recommendation weights (floating-point numbers accurate to one decimal place);" ": Semantic exploration weight (a floating-point number accurate to one decimal place).

[0062] Through the structured inputs described above, the large model can output based on clear logical principles. and And user intent. Furthermore, if the number of associated scenarios retrieved in the contextual memory layer is 0 (i.e., a completely new user or a completely new spatiotemporal scenario that has not been recorded), a degradation processing scheme is automatically triggered, setting the historical contextual memory in the prompt words to empty and setting the memory recommendation weight. Semantic exploration weight .

[0063] This example achieves accurate extraction and context construction of similar historical experiences through multi-dimensional scene feature overlap matching, effectively avoiding interference from irrelevant historical scenes on current decisions. By using structured prompts to constrain the boundaries of the large language model, it not only ensures the accuracy of intent parsing but also directly outputs controlled numerical proportions, transforming the originally difficult-to-quantify natural language context into clear quantitative recall allocation indicators, providing reliable data support for subsequent dual-path recommendation recall.

[0064] As an example, the multi-layered agent memory includes a working memory layer, a contextual memory layer, and a semantic memory layer. The step of filtering the content candidate pool based on global negative labels and scene negative labels extracted from the multi-layered agent memory to obtain a candidate list to be sorted includes: extracting global negative labels from the semantic memory layer; matching and eliminating candidate content in the content candidate pool based on the global negative labels to obtain a preliminary filtered content pool; extracting current scene data from the working memory layer and extracting scene negative labels corresponding to the current scene data from the contextual memory layer; matching and eliminating candidate content in the preliminary filtered content pool based on the scene negative labels to obtain a candidate list to be sorted.

[0065] It should be noted that in this example, the initial filtered content pool refers to the set of radio audios that remain after system-level content blocking. Matching and removal can be a data processing procedure that compares the unique identifiers and tag records of candidate content and removes the candidate content with a corresponding relationship from the current dataset.

[0066] Understandably, the first step is to perform global-level preference filtering. The backend system accesses the semantic memory layer and retrieves a global set of negative tags containing unique identifiers of audio files that users have long disliked. It then iterates through each candidate content in the content candidate pool, extracting its corresponding content ID. This extracted content ID is compared to the global set of negative tags; if a content ID exists in this set, the candidate content is directly removed from the content candidate pool. All the candidate content retained after this round of comparison constitutes the initial filtered content pool.

[0067] Next, fine-grained scene-level filtering is performed. Current scene data, including the time, location, and weather of the current session, is retrieved from the working memory layer. Using this current scene data as the retrieval criteria, a lookup table is performed in the context memory layer to extract a set of negative scene labels corresponding to the current environmental characteristics. This set specifically records unique audio identifiers that indicate a user's rejection behavior only in similar current scenarios.

[0068] Finally, a second round of filtering is performed on the initial filtered content pool based on negative scene tags. The candidate content in the initial filtered content pool is traversed, and the content ID is cross-referenced with the extracted set of negative scene tags. If a match is found, it means that although the candidate content is not globally blocked, it has generated negative feedback from users in the current environment, and therefore it is removed from the initial filtered content pool. After this round of matching and elimination, the final set of remaining audio files constitutes the candidate list to be sorted.

[0069] This example employs a dual filtering mechanism that separates global and scenario-based filtering, balancing basic user experience with contextual adaptability. Global filtering directly blocks content that users explicitly dislike, reducing the probability of compromised basic user experience. Scenario-based filtering precisely isolates inappropriate content in specific environments, improving the scenario relevance of radio content recommendations. This layered and progressive filtering approach not only increases the system's tolerance for negative feedback but also reduces the data computation scale in the subsequent weighted scoring stage, improving the overall data flow processing efficiency.

[0070] As an example, the step of performing multi-dimensional weighted scoring on candidate content in the unsorted candidate list based on the memory recommendation weight, the semantic exploration weight, and the user's immediate intent to obtain a recommended content sequence includes: when candidate content in the unsorted candidate list carries a memory tag, extracting the maximum scene score under the associated scene as a scene preference score, and setting the scene preference score as a preset initial score when it does not carry the memory tag; extracting the user's long-term interests from the semantic memory layer of the multi-layered agent memory library, and calculating the first similarity between the user's long-term interests and the content semantic features of the candidate content to obtain an interest preference score; calculating the second similarity between the user's immediate intent and the content semantic features to obtain an intent exploration score; performing a weighted summation of the scene preference score and the interest preference score based on the memory recommendation weight to obtain a memory path score; multiplying the semantic exploration weight by the intent exploration score to obtain an exploration path score, and summing the memory path score and the exploration path score to obtain a multi-dimensional weighted total score; and sorting the candidate content in descending order and truncating the position based on the multi-dimensional weighted total score to obtain a recommended content sequence.

[0071] It should be noted that in this example, the scenario preference score is a quantitative value used to measure the favorability of candidate content in similar historical environments. The preset initial score is a base score set for brand-new content that has not left any interaction records in the current associated scenario, and its value is 0. The semantic features of the content can be a high-dimensional vector representation that maps the text description or tags of the audio using a word embedding model. The first similarity and the second similarity are calculated indicators used to measure the semantic similarity between different feature vectors. The interest preference score is a quantitative score that reflects the degree of fit between the candidate content and the user's long-term auditory habits. The intent exploration score is a quantitative score that reflects the degree of fit between the candidate content and the user's direct needs in the current conversation. The memory path score and the exploration path score are stage-by-stage summary scores that represent the degree of continuity of historical preferences and the degree of fit of novel content, respectively. The multi-dimensional weighted total score is the final evaluation index derived by comprehensively considering historical environment performance, long-term inherent preferences, and current immediate needs. Position truncation is a filtering operation that extracts the top-ranked elements from an ordered set.

[0072] Understandably, the first step is to evaluate the historical performance of the candidate content within its context. Each candidate in the unsorted candidate list is iterated through, and each item is checked for a memory tag. If a memory tag is present, it indicates that the content originated from recalls in similar historical environments. Multiple scene scores accumulated by the content across various associated scenarios are extracted, compared, and the maximum score is selected as the scene preference score. If no memory tag is present, it indicates that the content belongs to unfamiliar audio retrieved based on intent exploration and lacks historical scene verification. Therefore, the scene preference score is directly assigned to a preset initial score.

[0073] Next, the long-term auditory preference matching degree is calculated. The recorded long-term user interest text is read from the semantic memory layer and converted into a long-term interest feature vector; the description or tag text data of candidate content is obtained and converted into content semantic features using a text embedding model. The cosine similarity between the long-term interest feature vector and the content semantic features is calculated, and the resulting cosine similarity is used as the first similarity score, which is then used as the interest preference score.

[0074] Next, the short-term immediate appeal matching degree is calculated. The user's immediate intent text parsed from the context is obtained and converted into an immediate intent feature vector in the same way. The cosine similarity between the immediate intent feature vector and the content semantic features is calculated, and the resulting cosine similarity is used as the second similarity. This value is used as the intent exploration score.

[0075] Furthermore, the calculation metrics from various dimensions are integrated. The scene preference score and interest preference score are added together, and then multiplied by a pre-assigned memory recommendation weight to obtain the memory path score representing the familiarity dimension. The intent exploration score is directly multiplied by a pre-assigned semantic exploration weight to obtain the exploration path score representing the novelty dimension. The memory path score and exploration path score are added together to obtain a multi-dimensional weighted total score reflecting the overall recommendation value of the candidate content. The formula for calculating the multi-dimensional weighted total score is as follows: in, This refers to the multi-dimensional weighted total score of candidate content i; This refers to the weight of memory recommendations; This refers to semantic exploration weights; It refers to the scene integral adjustment coefficient, which is used to scale the absolute value of the scene integral to fit the sensitive range of the Sigmoid function; This refers to the scenario preference score; It refers to the time decay factor, used to calculate the timeliness of user scenario preferences; It refers to the time difference, which is the difference between the last time the content was interacted with and the current time of interaction; This refers to the preset long-term interest weighting coefficient; This refers to interest preference scores; This refers to the score for intention exploration; This refers to the Sigmoid function. Because the numerical dimensions of historically accumulated scene preference scores and cosine similarity are inconsistent, the Sigmoid function is used instead. Cooperate The data is nonlinearly normalized to converge to the (0,1) interval so that it can be linearly weighted with the interest preference score.

[0076] Finally, sorting and list generation are performed. All candidate content in the candidate list to be sorted is sorted in descending order according to the weighted total score. After sorting, position truncation is performed, retaining a predetermined number of candidate audios with the highest ranking (e.g., the top 10), thus forming the final recommended content sequence for playback.

[0077] This example effectively integrates historical contextual experience, long-standing habits, and current immediate needs by constructing a multi-dimensional scoring system. Different starting scores are applied to content from different recall paths, ensuring the objectivity and rationality of the evaluation process. By calculating the stage scores for memory paths and exploration paths separately and using dynamic weights for overall consideration, a smooth transition between historical continuity and unknown exploration of recommended resources is achieved. The final generated sequence not only ensures the familiarity of contextual companionship but also expands with novel content in a timely manner while meeting immediate intentions, effectively improving the overall experience of audio recommendation.

[0078] As an example, the step of generating an interactive introduction based on the source tag, the memory recommendation weight, and the semantic exploration weight, and combining the interactive introduction with the audio corresponding to the recommended content sequence to obtain radio content includes: matching a corresponding introduction generation strategy in a preset strategy library based on the source tag, the memory recommendation weight, and the semantic exploration weight; determining an emotion type vector and an emotion intensity coefficient based on the introduction generation strategy; extracting attribute information from the recommended content sequence, and constructing a prompt word template using the attribute information, the user's immediate intent, and the current scene data in the working memory layer of the multi-layer intelligent agent memory library; inputting the prompt word template into the large language model for text generation to obtain interactive introduction text; converting the interactive introduction text into an introduction speech stream based on the emotion type vector and the emotion intensity coefficient; and concatenating and outputting the introduction speech stream with the audio corresponding to the recommended content sequence to obtain radio content.

[0079] It should be noted that in this example, the preset strategy library refers to a pre-configured mapping set containing various introductory content generation rules and corresponding acoustic sentiment parameters. The introductory generation strategy refers to a predefined framework based on the audio recommendation source and user intent, used to guide text generation tendencies and the tone of speech synthesis. The sentiment type vector can be a control parameter used to instruct the speech synthesis model to output a specific core sentiment type (e.g., warm, gentle, or lively). The sentiment intensity coefficient is a continuous numerical variable used to adjust the tension and fluctuation of core emotional expression. Attribute information can be basic descriptive metadata such as the song title, artist, audio genre, or album name of the recommended audio. The cue word template refers to a text structure assembled from discrete context data according to a specific logical framework, used to guide the large language model to generate compliant interactive statements. The interactive introductory text refers to a natural language segment generated by the model and used to be broadcast to the user before the formal audio playback. The introductory speech stream refers to an audio waveform sequence with emotional features generated by converting natural language text through an acoustic model. The radio content can be a continuous data stream ultimately presented to the user, containing a mixture of broadcast segments and media audio.

[0080] Understandably, the first step is to determine the appropriate broadcast strategy and emotional tone. This involves reading the source tag, memory recommendation weight, and semantic exploration weight of the top-ranked recommended content. Conditional matching is then performed within a pre-defined strategy library: when the source tag is a memory tag and the memory recommendation weight is significantly greater than the semantic exploration weight, the contextual resonance strategy is matched, extracting the corresponding emotional type vector (e.g., a feature parameter representing "warmth") and setting the emotional intensity coefficient to 1.0; when the source tag is a memory tag and the semantic exploration weight is not less than the memory recommendation weight, the tacit response strategy is matched, extracting the corresponding emotional type vector (e.g., a feature parameter representing "gentleness") and setting the emotional intensity coefficient to 0.9; when the source tag is a semantic tag, the semantic exploration strategy is directly matched, extracting the corresponding emotional type vector (e.g., a feature parameter representing "lively") and setting the emotional intensity coefficient to 1.1.

[0081] Next, the input context of the large language model is assembled. Attribute information of the recommended content is extracted from the media library. Simultaneously, the previously parsed user's immediate intent text is retrieved, and current scene data (such as current weather conditions and time period) is extracted from the working memory layer. The acquired attribute information, user's immediate intent, current scene data, and matched lead generation strategy requirements are then filled into the prompt word template according to the system-configured paragraph structure, forming a complete system-level text instruction.

[0082] Furthermore, interactive speech is generated and synthesized. The constructed prompt word template is input into a large language model, which outputs an interactive introductory text that fits the current context based on the environmental state and recommendation reasons in the template. Subsequently, a text-to-speech (TTS) model is invoked, using the interactive introductory text as the basic conversion object, while simultaneously inputting the previously determined sentiment type vector and sentiment intensity coefficient as acoustic adjustment parameters. The TTS model combines these parameters to control the pitch, speech rate, and pauses during the synthesis process, outputting an introductory speech stream that conforms to the corresponding policy's emotional tone.

[0083] Finally, audio sequence splicing is performed. The generated introductory audio stream is placed before the recommended content audio file in the server-side or terminal player cache, spliced ​​and merged in chronological order, and output as continuous, seamless radio content for playback.

[0084] In this example, the system switches the introductory prompt template based on the source tags and weights output in the preceding sequence to ensure that the text style, recommendation logic and emotional tone of the introductory text are highly aligned. To achieve fine voice control, the system introduces an emotion vector (emo_vector) to specify the core emotion type and introduces an emotion intensity coefficient (emo_alpha) to control the expressive tension of the emotion. The specific generation strategy and parameter configuration are as follows: (1) When the source is a memory tag and The scenario resonance strategy is triggered at the time. This strategy aims to convey a tone of nostalgia and companionship. The emotional vector emo_vector is set to "warm", and the emotional intensity coefficient emo_alpha is set to 1.0. Through the warm emotional tension, the user is evoked with a sense of companionship in a familiar scene. (2) When the source is a memory tag and When the tacit response strategy is triggered, although the recommended content is memory-based, it is highly consistent with the user's immediate intent. The aim is to convey the tone of tacit response, with the emotion vector emo_vector set to "gentle" and the emotion intensity coefficient emo_alpha set to 0.9. It is like an intelligent assistant that understands the user explaining the reasons for the recommendation. (3) When the source is a semantic tag, the semantic exploration strategy is triggered to transparently explain the objective reasons for recommending the content to the user. The aim is to convey the tone of intent consistency and exploration, with the emotion vector emo_vector set to "lively" and the emotion intensity coefficient emo_alpha set to 1.1. The lively background is overlaid with a moderately rising tone, which effectively arouses the user's attention and desire to explore.

[0085] This example dynamically switches the introductory text generation strategy based on different recommendation tracing paths and weight ratios, enabling the system to output interactive voice messages with differentiated emotional nuances for different scenarios such as "nostalgic companionship," "favorable recommendations," or "novel exploration." By utilizing a large language model to integrate real-time context and recommendation reasons, and supplementing it with speech synthesis technology with emotional parameter adjustment, the rigidity and factual misalignment caused by traditional preset fixed broadcast templates are eliminated. The overall process achieves a natural interactive closed loop, ensuring a high degree of alignment between the introductory text content, the emotional expression of the voice, and the actual played audio, significantly enhancing the anthropomorphic companionship experience of the smart radio.

[0086] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 , Figure 2 This is a flowchart illustrating the second embodiment of the radio content generation method based on spatiotemporal scene association and agent memory of this application. Step S30 of the radio content generation method based on spatiotemporal scene association and agent memory includes steps S31 to S36: Step S31: Determine the first recall quantity corresponding to the memory path according to the memory recommendation weight, and retrieve the first candidate content with a scene score greater than or equal to a preset score threshold in the radio content library according to the first recall quantity. Step S32: Attach memory tags to the first candidate content to obtain memory-oriented recall results; Step S33: Determine the second recall quantity corresponding to the exploration path based on the semantic exploration weight, and convert the user's immediate intent into an intent vector; Step S34: In the radio content library, based on the second recall quantity, recall the second candidate content that matches the intent vector using an approximate nearest neighbor retrieval algorithm; Step S35: Add semantic tags to the second candidate content to obtain semantic exploration recall results; Step S36: The memory-oriented recall results and the semantic exploration recall results are deduplicated and merged to obtain a content candidate pool with source tags.

[0087] It should be noted that in this example, the first recall quantity refers to the specific number of audio samples extracted from the total capacity of the candidate pool based on historical preferences. The second recall quantity refers to the specific number of audio samples extracted from the global media library based on novel or unknown audio samples. The preset score threshold is the minimum score threshold for determining whether historical audio has recommendation value, set to 60, representing that the corresponding content's preference level has reached the passing standard. The first candidate content can be a set of high-scoring historical audio samples that meet the memory path recall conditions. The second candidate content can be a set of entirely new audio samples that meet the exploration path recall conditions. Memory tags refer to specific attribute identifiers that represent the source of audio recommendations as a continuation of historical preferences. Semantic tags refer to specific attribute identifiers that represent the source of audio recommendations as a new intent exploration. Intent vectors refer to multi-dimensional numerical arrays transformed from user requests expressed in natural language through word embedding technology. Approximate nearest neighbor retrieval algorithms can be retrieval algorithms that quickly find the closest set of vectors in high-dimensional space based on mechanisms such as Locality Sensitive Hash (LSH) or Hierarchical Navigation Small World Graph (HNSW).

[0088] Understandably, the process begins with executing a recall operation corresponding to the memory path. The system's pre-defined total candidate pool target capacity is obtained. This target capacity is then multiplied by the memory recommendation weight to calculate the first recall quantity. Next, in the radio content library, audio clips with scene scores greater than or equal to a preset score threshold (60) are retrieved based on historical interaction records. These clips are then extracted in descending order of scene scores, retaining the top-ranked first recall quantity of historical audio clips to form the first candidate content. Subsequently, a memory tag is attached to each audio clip in the first candidate content, using this tag to represent the continuation of historical preferences. After tagging, the memory-oriented recall result is obtained.

[0089] Then, the recall operation corresponding to the exploration path is executed. The total candidate pool target capacity is multiplied by the semantic exploration weight to calculate the second recall quantity. The user's real-time intent text is extracted and converted into a high-dimensional intent vector by a pre-trained language model such as a text embedding model. A global audio feature vector is generated in advance by converting the audio's attribute tags and descriptive information using a text embedding model. Then, in the radio content library, the spatial distance between the intent vector and the global audio feature vector is calculated using an approximate nearest neighbor retrieval algorithm. The top two recall quantity audios with the closest spatial distance are extracted to form the second candidate content. Next, a semantic tag is attached to each audio in the second candidate content. This tag is used to represent the new intent exploration, and the semantic exploration recall result is obtained after the tagging is completed.

[0090] Finally, the data from the dual-path recall is integrated. The results of memory-oriented recall and semantic exploration recall are merged. During the merging process, unique identifiers of the audio are extracted and cross-referenced. If the same identifier is found to exist in both sets of results, the record with the memory tag is retained, and the other duplicate is removed, completing the deduplication process. After merging and deduplication, a content candidate pool with source tags is obtained.

[0091] This embodiment achieves a reasonable allocation of resources between maintaining users' existing auditory habits and exploring potential new interests by splitting the recall process into two independent paths: memory orientation and semantic exploration, and using dynamic weights to calculate the recall quantity of each. The approximate nearest neighbor retrieval algorithm significantly improves the vector matching efficiency in the massive media library. By attaching special source tags to content from different sources, it not only clearly records the generation logic of recommended resources, but also provides key judgment criteria for subsequent multi-dimensional weighted scoring and differentiated introduction generation, ensuring the interpretability and coherence of the content distribution process.

[0092] For example, to help understand the implementation process of the radio content generation method based on spatiotemporal scene association and agent memory obtained by combining this embodiment with the above embodiment one, please refer to Figure 3 , Figure 3 A simplified flowchart is provided for a radio content generation method based on spatiotemporal scene association and agent memory. Specifically: First, in the data acquisition phase, the system acquires multimodal spatiotemporal environmental perception data including time, weather, and location; user interaction behavior data including interactive text, completion playback, and song switching information; and data from the radio content library. This data flows down to the agent memory and the memory update and maintenance module. The agent memory includes a working memory layer storing the current dialogue context and multimodal scene data; a contextual memory layer storing historical scene dialogue summaries and interaction content; and a semantic memory layer storing long-term user interests and global and scene negative labels. Simultaneously, the memory update and maintenance module triggers multi-layered memory updates at the end of the dialogue, using a large model to summarize the dialogue summary and user interests, and performing points management and scene content cooldown / revival operations. Next, the process enters the memory retrieval enhanced large model intent parsing and decision-making module, sequentially going through scene data matching, retrieving historical dialogue summaries in similar scenes, and retrieval-enhanced LLM inference steps, ultimately outputting weights and user intent.

[0093] The system then performs two-way recall based on the output results. First, it uses memory-oriented recall based on scene data matching to recall high-scoring content and assign it a memory tag. Second, it uses semantic exploration recall to convert user intent into a vector for retrieval to recall related content and assign it a semantic tag. The results of these two recalls are combined into a candidate content pool. The content in the candidate pool then enters a dual content filtering module, undergoing memory filtering to filter globally negative recall content and scene filtering to filter scene-specific negative recall content. After filtering, the content enters a weighted ranking calculation stage that integrates memory preference and intent exploration weights. The system calculates memory preference-related scores and intent exploration-related scores respectively, and combines them with the formula in the diagram to derive the total ranking score. The ranked and selected content then enters a multi-strategy interactive introductory text generation module, sequentially performing multi-strategy-driven introductory text generation instruction matching, strategy-, memory-, and recommended content-based prompt word construction, and LLM introductory text generation. Finally, in the TTS speech conversion and audio output module, TTS emotional introductory audio is generated, and the generated introductory text is concatenated with the recommended audio content to complete the output process of playing the synthesized audio.

[0094] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the radio content generation method based on spatiotemporal scene association and agent memory of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0095] This application also provides a radio content generation device based on spatiotemporal scene association and agent memory. Please refer to [reference needed]. Figure 4 The radio content generation device based on spatiotemporal scene association and agent memory includes: The memory update module 10 is used to update the multi-layer intelligent agent memory bank based on external spatiotemporal characteristics and user interaction data. The intent parsing module 20 is used to parse the intent of the multi-layer intelligent agent memory database through a large language model to obtain the user's immediate intent, memory recommendation weight, and semantic exploration weight. The dual-path recall module 30 is used to perform memory-oriented recall based on the memory recommendation weight and semantic exploration recall based on the semantic exploration weight, so as to obtain a content candidate pool with source tracing tags. The negative filtering module 40 is used to filter the content candidate pool based on the global negative labels and scene negative labels extracted from the multi-layer intelligent agent memory library to obtain a candidate list to be sorted. The multi-dimensional scoring module 50 is used to perform multi-dimensional weighted scoring on the candidate content in the candidate list to be sorted based on the memory recommendation weight, the semantic exploration weight, and the user's real-time intent, so as to obtain a recommended content sequence. The content synthesis module 60 is used to generate an interactive introduction based on the source tag, the memory recommendation weight, and the semantic exploration weight, and to combine the interactive introduction with the audio corresponding to the recommended content sequence to obtain radio content.

[0096] The radio content generation device based on spatiotemporal scene association and agent memory provided in this application, employing the radio content generation method based on spatiotemporal scene association and agent memory in the above embodiments, can solve the technical problem of how to enhance the personalized companionship of smart radio and provide users with an immersive listening experience that is coherent across scenes and emotionally natural. Compared with the prior art, the beneficial effects of the radio content generation device based on spatiotemporal scene association and agent memory provided in this application are the same as the beneficial effects of the radio content generation method based on spatiotemporal scene association and agent memory provided in the above embodiments, and other technical features in the radio content generation device based on spatiotemporal scene association and agent memory are the same as the features disclosed in the methods of the above embodiments, and will not be repeated here.

[0097] This application provides a radio content generation device based on spatiotemporal scene association and agent memory. The radio content generation device based on spatiotemporal scene association and agent memory includes: at least one processor; and a memory communicatively connected to at least one processor; wherein the memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor to enable at least one processor to execute the radio content generation method based on spatiotemporal scene association and agent memory in the above embodiment 1.

[0098] The following is for reference. Figure 5 This document illustrates a structural schematic diagram of a radio content generation device based on spatiotemporal scene association and agent memory, suitable for implementing embodiments of this application. The radio content generation device based on spatiotemporal scene association and agent memory in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Android Devices), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 5 The radio content generation device based on spatiotemporal scene association and agent memory shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0099] like Figure 5As shown, the radio content generation device based on spatiotemporal scene association and agent memory may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to programs stored in ROM 1002 (Read Only Memory) or programs loaded from storage device 1003 into RAM 1004 (Random Access Memory). RAM 1004 also stores various programs and data required for the operation of the radio content generation device based on spatiotemporal scene association and agent memory. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. I / O interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, LCDs (Liquid Crystal Displays), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the radio content generation device based on spatiotemporal scene association and agent memory to exchange data with other devices wirelessly or via wired communication. Although the figure shows a radio content generation device based on spatiotemporal scene association and agent memory with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0100] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0101] The radio content generation device based on spatiotemporal scene association and intelligent agent memory provided in this application, employing the radio content generation method based on spatiotemporal scene association and intelligent agent memory in the above embodiments, can solve the technical problem of how to enhance the personalized companionship of intelligent radio and provide users with a cross-scene coherent and emotionally natural immersive listening experience. Compared with the prior art, the beneficial effects of the radio content generation device based on spatiotemporal scene association and intelligent agent memory provided in this application are the same as the beneficial effects of the radio content generation method based on spatiotemporal scene association and intelligent agent memory provided in the above embodiments, and other technical features in this radio content generation device based on spatiotemporal scene association and intelligent agent memory are the same as the features disclosed in the previous embodiment method, and will not be repeated here.

[0102] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0103] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0104] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the radio content generation method based on spatiotemporal scene association and agent memory in the above embodiments.

[0105] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or Flash Memory), optical fibers, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0106] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by a radio content generation device based on spatiotemporal scene association and agent memory, the device performs the following actions: updates a multi-layer agent memory database based on external spatiotemporal features and user interaction data; performs intent parsing on the multi-layer agent memory database using a large language model to obtain the user's immediate intent, memory recommendation weight, and semantic exploration weight; performs memory-oriented recall based on the memory recommendation weight and semantic exploration recall based on the semantic exploration weight to obtain a content candidate pool with source tags; filters the content candidate pool based on global negative tags and scene negative tags extracted from the multi-layer agent memory database to obtain a candidate list to be sorted; performs multi-dimensional weighted scoring on the candidate content in the candidate list to be sorted based on the memory recommendation weight, the semantic exploration weight, and the user's immediate intent to obtain a recommended content sequence; generates an interactive introduction based on the source tags, the memory recommendation weight, and the semantic exploration weight, and combines the interactive introduction with the audio corresponding to the recommended content sequence to obtain radio content.

[0107] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LAN (Local Area Network) or WAN (Wide Area Network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0108] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0109] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the radio content generation method based on spatiotemporal scene association and agent memory described above. This addresses the technical problem of how to enhance the personalized companionship of smart radio stations and provide users with a cross-scene, coherent, and emotionally natural immersive listening experience. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the radio content generation method based on spatiotemporal scene association and agent memory provided in the above embodiments, and will not be elaborated upon here.

[0110] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the radio content generation method based on spatiotemporal scene association and agent memory as described above.

[0111] The computer program product provided in this application solves the technical problem of how to enhance the personalized companionship of smart radio and provide users with an immersive listening experience that is coherent across scenarios and emotionally natural. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the radio content generation method based on spatiotemporal scene association and agent memory provided in the above embodiments, and will not be repeated here.

[0112] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for generating radio content based on spatiotemporal scene association and agent memory, characterized in that, The method includes: The multi-layer intelligent agent memory bank is updated based on external spatiotemporal characteristics and user interaction data. The multi-layer intelligent agent memory bank refers to a hierarchical storage architecture consisting of a working memory layer, a contextual memory layer, and a semantic memory layer, which are used for short-term state recording, historical scene snapshot storage, and long-term interest extraction, respectively. The intent is parsed by the multi-layer intelligent agent memory database through a large language model to obtain the user's immediate intent, memory recommendation weight, and semantic exploration weight. The memory recommendation weight refers to the proportion of content that the user has historically preferred when allocating recall resources, and its value ranges from 0.0 to 1.

0. The semantic exploration weight is the proportion of content that is allocated to new or untouched content, and its value ranges from 0.0 to 1.

0. The sum of the semantic exploration weight and the memory recommendation weight is strictly equal to 1.

0. Based on the memory recommendation weight, memory-oriented recall is performed, and based on the semantic exploration weight, semantic exploration recall is performed to obtain a content candidate pool with source tracing tags; The content candidate pool is filtered based on the global negative labels and scene negative labels extracted from the multi-layer intelligent agent memory library to obtain a candidate list to be sorted. Based on the memory recommendation weight, the semantic exploration weight, and the user's immediate intent, the candidate content in the unsorted candidate list is scored using a multi-dimensional weighted score to obtain a recommended content sequence. An interactive introduction is generated based on the source tag, the memory recommendation weight, and the semantic exploration weight. The interactive introduction is then combined with the audio corresponding to the recommended content sequence to obtain radio content.

2. The method as described in claim 1, characterized in that, The multi-layered intelligent agent memory bank includes a working memory layer, a contextual memory layer, and a semantic memory layer; The step of updating the multi-layer intelligent agent memory bank based on external spatiotemporal characteristics and user interaction data includes: The external spatiotemporal features are discretized to obtain the current scene data, and the original dialogue text sequence and listening behavior feedback are parsed from the user interaction data. The current scene data and the original dialogue text sequence are stored in the working memory layer; Update the scene score of the corresponding scene in the scene memory layer based on the listening behavior feedback; When the scene integral is lower than a preset negative threshold, a scene negative label and a global negative label are generated; The original dialogue text sequence is semantically compressed using a large language model to generate a dialogue summary, and features are extracted from the dialogue summary using the large language model to obtain the user's long-term interests. The scene negative label, the dialogue summary, and the scene score are stored in the context memory layer, and the global negative label and the user's long-term interests are stored in the semantic memory layer, thus completing the update.

3. The method as described in claim 2, characterized in that, The step of parsing the multi-layer agent memory database using a large language model to obtain the user's immediate intent, memory recommendation weight, and semantic exploration weight includes: Retrieve related scenes from the scene memory layer that have a feature overlap of greater than or equal to a preset overlap threshold with the current scene data in a preset dimension; Extract the dialogue summary corresponding to the associated scenario, and concatenate the dialogue summary into historical context according to the time sequence; The preset weighted inference rules, the historical context, the current scene data, and the original dialogue text sequence are used to construct structured prompt words; The structured prompts are input into the large language model for quantitative reasoning calculations to obtain the user's immediate intent, memory recommendation weight, and semantic exploration weight.

4. The method as described in claim 1, characterized in that, The steps of performing memory-oriented recall based on the memory recommendation weight and semantic exploration recall based on the semantic exploration weight to obtain a content candidate pool with source tracing tags include: The first recall quantity corresponding to the memory path is determined based on the memory recommendation weight, and the first candidate content with a scene score greater than or equal to a preset score threshold is retrieved from the radio content library based on the first recall quantity. Attach memory tags to the first candidate content to obtain memory-oriented recall results; The second recall quantity corresponding to the exploration path is determined based on the semantic exploration weight, and the user's immediate intent is converted into an intent vector; In the radio content library, based on the second recall quantity, a second candidate content matching the intent vector is recalled using an approximate nearest neighbor retrieval algorithm; Add semantic tags to the second candidate content to obtain semantic exploration recall results; The memory-oriented recall results and the semantic exploration recall results are deduplicated and merged to obtain a content candidate pool with source tags.

5. The method as described in claim 1, characterized in that, The multi-layered intelligent agent memory bank includes a working memory layer, a contextual memory layer, and a semantic memory layer; The step of filtering the content candidate pool based on global negative labels and scene negative labels extracted from the multi-layer agent memory to obtain a candidate list to be sorted includes: Extract global negative labels from the semantic memory layer; Based on the global negative tags, candidate content in the content candidate pool is matched and eliminated to obtain a preliminary filtered content pool; Extract current scene data from the working memory layer, and extract scene negative labels corresponding to the current scene data from the context memory layer; Based on the negative tags of the scenario, candidate content in the initial filtered content pool is matched and eliminated to obtain a candidate list to be sorted.

6. The method as described in claim 1, characterized in that, The step of performing multi-dimensional weighted scoring on the candidate content in the candidate list to be sorted based on the memory recommendation weight, the semantic exploration weight, and the user's immediate intent to obtain the recommended content sequence includes: When the candidate content in the unsorted candidate list carries a memory tag, the maximum scene score under the associated scene is extracted as the scene preference score, and when the memory tag is not carried, the scene preference score is set to a preset initial score. The user's long-term interests are extracted from the semantic memory layer of the multi-layer intelligent agent memory library, and the first similarity between the user's long-term interests and the content semantic features of the candidate content is calculated to obtain the interest preference score. Calculate the second similarity between the user's immediate intent and the semantic features of the content to obtain the intent exploration score; The memory path score is obtained by weighting and summing the scene preference score and the interest preference score based on the memory recommendation weight. Multiply the semantic exploration weight by the intent exploration score to obtain the exploration path score, and sum the memory path score with the exploration path score to obtain the multidimensional weighted total score; Based on the multidimensional weighted total score, the candidate content is sorted in descending order and its position is truncated to obtain the recommended content sequence.

7. The method according to any one of claims 1 to 6, characterized in that, The step of generating an interactive introduction based on the source tag, the memory recommendation weight, and the semantic exploration weight, and combining the interactive introduction with the audio corresponding to the recommended content sequence to obtain the radio content includes: Based on the source tag, the memory recommendation weight, and the semantic exploration weight, a corresponding lead generation strategy is matched in the preset strategy library; The sentiment type vector and sentiment intensity coefficient are determined based on the aforementioned lead generation strategy; Extract the attribute information of the recommended content sequence, and construct the prompt word template by combining the attribute information, the user's immediate intent, and the current scene data in the working memory layer of the multi-layer intelligent agent memory bank; The prompt word template is input into the large language model to generate text, resulting in interactive introductory text. Based on the emotion type vector and the emotion intensity coefficient, the interactive introductory text is converted into an introductory speech stream; The introductory audio stream is concatenated with the audio corresponding to the recommended content sequence and then output to obtain the radio content.

8. A radio content generation device based on spatiotemporal scene association and agent memory, characterized in that, The device employs the radio content generation method based on spatiotemporal scene association and agent memory as described in any one of claims 1 to 7, and the device comprises: The memory update module is used to update the multi-layer intelligent agent memory bank according to external spatiotemporal characteristics and user interaction data. The multi-layer intelligent agent memory bank refers to a hierarchical storage architecture consisting of a working memory layer, a contextual memory layer, and a semantic memory layer, which are used for short-term state recording, historical scene snapshot storage, and long-term interest extraction, respectively. The intent parsing module is used to parse the intent of the multi-layer intelligent agent memory bank through a large language model to obtain the user's immediate intent, memory recommendation weight, and semantic exploration weight. The memory recommendation weight refers to the proportion of historically preferred content assigned to the user when allocating recall resources, with a value range of 0.0 to 1.

0. The semantic exploration weight is the proportion of new or untouched content assigned to the user, with a value range of 0.0 to 1.

0. The sum of the semantic exploration weight and the memory recommendation weight is strictly equal to 1.

0. The dual-path recall module is used to perform memory-oriented recall based on the memory recommendation weight and semantic exploration recall based on the semantic exploration weight, so as to obtain a content candidate pool with source tracing tags. The negative filtering module is used to filter the content candidate pool based on global negative labels and scene negative labels extracted from the multi-layer intelligent agent memory library to obtain a candidate list to be sorted. The multi-dimensional scoring module is used to perform multi-dimensional weighted scoring on the candidate content in the candidate list to be sorted based on the memory recommendation weight, the semantic exploration weight, and the user's real-time intent, so as to obtain a recommended content sequence. The content synthesis module is used to generate an interactive introduction based on the source tag, the memory recommendation weight, and the semantic exploration weight, and to combine the interactive introduction with the audio corresponding to the recommended content sequence to obtain radio content.

9. A radio content generation device based on spatiotemporal scene association and intelligent agent memory, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the radio content generation method based on spatiotemporal scene association and agent memory as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the radio content generation method based on spatiotemporal scene association and agent memory as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Information recommendation method and system based on multi-path recall

    CN115309996A

  • Intelligent radio content dynamic adjustment method based on scene perception

    CN121117254A