Methods, apparatus, and devices for optimizing generation efficiency based on intent recognition and memory
By using a pre-trained intent recognition and analysis model and a memory augmentation network, the generation mode is dynamically selected and the network weights are optimized, which solves the problems of response latency and wasted computing resources in AIGC devices, and improves generation efficiency and content quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NINGBO SIMSHINE INTELLIGENT TECH CO LTD
- Filing Date
- 2026-03-10
- Publication Date
- 2026-06-02
AI Technical Summary
Existing AIGC devices have shortcomings in response latency, computational resource consumption, context continuity, and user experience. In particular, they fail to effectively combine the differences in the complexity of user requests with the lightweight requirements of device usage scenarios, resulting in wasted computational resources and untimely responses.
By using a pre-trained intent recognition and analysis model, the urgency and complexity level of user audio data are quantified, the generation mode is dynamically selected, and a memory enhancement network is used to obtain matching memory feature information to generate target content. The network weights are adjusted based on user feedback to optimize generation efficiency.
It achieves precise matching of generation strategies, improves generation response efficiency, reduces computing resource consumption, and makes the generated results more in line with user needs, thereby improving the quality and consistency of the generated content.
Smart Images

Figure CN122135718A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and in particular to a method, apparatus, and device for optimizing generation efficiency based on intent recognition and memory. Background Technology
[0002] With the rapid development of AI-generated content (AIGC) technology, smart interactive hardware is increasingly being used in various fields, especially in children's education and creative design. Devices such as AI sticker printers and interactive learning machines can generate personalized images, text, or other content based on user-input audio commands or text descriptions. These devices rely on powerful generative models to efficiently transform user intentions into concrete visual representations or creative content.
[0003] However, existing AIGC devices generally face several challenges, particularly in terms of response latency, computational resource consumption, context continuity, and user experience. For example, regardless of whether the user inputs a simple request or a complex creative description, the default approach is to invoke a large-scale cloud-based generation model for complete, in-depth reasoning, without considering the complexity of user requests or the lightweight requirements of the device's usage scenario. This wastes computational resources and fails to respond quickly to user needs. Existing patent CN120893588A (Intent-Aware Dynamic Selection and Optimization Method, System, and Device for Retrieval Path) can segment main / sub-intents, identify causal relationships, and select the optimal retrieval path based on these relationships. However, it cannot dynamically select a generation strategy based on the user's input intent, nor can it achieve context continuity and personalized learning of user preferences. Furthermore, improvements are still needed in the quality and consistency of the generated content.
[0004] There is an urgent need for a new technological approach to address the response latency issue in the AI-generated content process and improve the user experience of AIGC devices. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a method, apparatus and device for optimizing generation efficiency based on intent recognition and memory, in order to solve the response delay problem in the process of artificial intelligence generating content in the prior art.
[0006] In a first aspect, embodiments of the present invention provide a generation efficiency optimization method based on intent recognition and memory, the method comprising: The raw audio data is input into a pre-trained intent recognition and analysis model to obtain text information, urgency level, and content complexity level corresponding to the raw audio data. Based on the urgency level and content complexity level, determine the target generation mode; Memory feature information matching the text information is obtained through a preset memory enhancement network; Based on the text information, memory feature information, and target generation pattern, generate target-generated content; Obtain user feedback on the target generative content; Based on the feedback information, the weights of the memory enhancement network are adjusted to optimize the generation efficiency.
[0007] Preferably, the step of inputting the original audio data into a pre-trained intent recognition and analysis model to obtain text information, urgency level, and content complexity level corresponding to the original audio data includes: The original audio data is preprocessed to obtain the target audio data; Based on the audio recognition algorithm, the target audio data is subjected to text recognition to obtain text information; Natural language understanding is performed on the text information to obtain a semantic feature vector, wherein the semantic feature vector is used to represent the number of entities, the structural relationship features between entities, and the number of unconventional adjectives; According to the preset urgency rating criteria, the target audio data is rated for urgency to obtain an urgency level, wherein the urgency level includes high urgency and low urgency. Fuzzy reasoning is performed on the semantic feature vector to obtain the content complexity level, wherein the content complexity level includes low complexity, medium complexity and high complexity.
[0008] Preferably, the step of rating the target audio data for urgency according to a preset urgency rating standard to obtain an urgency level includes: Based on preset dual thresholds, endpoint detection is performed on the target audio data to obtain valid audio segments; Syllable identification and counting are performed on the effective syllable segments to obtain the number of syllables within the effective syllable segments; The real-time speech rate is calculated based on the ratio of the number of syllables to the effective speech duration. Based on the preset speech rate emergency assessment rules, the real-time speech rate, and the pre-collected user-personalized baseline speech rate, the speech rate emergency characteristics are obtained. The short-time energy value of each frame of the target audio data is calculated based on the short-time energy calculation algorithm. The short-time energy value is weighted, filtered, and logarithmically transformed to obtain the real-time volume. Based on the real-time volume and the preset volume evaluation rules, the volume emergency characteristics are obtained; Based on the preset emergency keyword database and the text information, the emergency characteristics of the keywords are obtained; An urgency level is obtained based on the urgency characteristics of speech rate and volume, as well as the preset urgency rating criteria.
[0009] Preferably, the step of performing fuzzy reasoning on the semantic feature vector to obtain the content complexity level includes: The number of entities, the structural association features between entities, and the number of unconventional adjectives in the semantic feature vector are normalized respectively to obtain each confidence score; Based on the preset membership function and the preset complexity scoring rule library, fuzzy reasoning and defuzzification are performed on each confidence score to obtain a continuous complexity score. The content complexity level is obtained based on the continuous complexity score and the preset dynamic complexity decision threshold.
[0010] Preferably, determining the target generation mode based on the urgency level and content complexity level includes: When the content complexity level is low and the urgency level is high, the target generation mode is determined to be the rapid generation mode. The rapid generation mode refers to the generation mode that directly calls the local template library without calling any model. When the content complexity level is low and the urgency level is low, the target generation mode is determined to be a balanced generation mode, wherein the balanced generation mode refers to the generation mode that calls the local lightweight model. When the content complexity level is medium complexity and the urgency level is low urgency, the target generation mode is determined to be a balanced generation mode. When the content complexity level is medium complexity and the urgency level is high urgency, the target generation mode is determined to be the rapid generation mode. When the content complexity level is high complexity, the target generation mode is determined to be a refined generation mode, wherein the refined generation mode refers to the generation mode using a large cloud model; When the content complexity level is high complexity and the urgency level is high urgency, the target generation mode is determined to be the balanced generation mode.
[0011] Preferably, obtaining user feedback information on the target generative content includes: Acquire user behavior feedback data on target generated content, wherein the behavior feedback data includes gaze duration, operation behavior data, head posture features, repetition request interval and touch pressure data; Based on preset judgment rules and behavioral feedback data, the user's feedback type and feedback score are obtained, wherein the feedback type includes positive feedback and negative feedback; Based on the feedback score and feedback type, the user's feedback information on the target generative content is obtained.
[0012] Preferably, the step of obtaining user behavioral feedback data on the target generated content includes: Acquire several consecutive frames of images from the user; Input several of the images into a preset human posture detection model to obtain the head yaw angle, head pitch angle and coordinates of key human points; Based on the difference in head yaw angle and head pitch angle between adjacent frames and a preset angle difference threshold, it is determined whether the user is looking at the generated content. When it is determined that a user is gazing at the generative content, the gazing duration is obtained; Based on the coordinates of the key human body points and the preset posture judgment template, the head posture features are obtained; Obtain a first timestamp of the target generated content display and a second timestamp of the user operation, wherein the user operation includes skipping, saving, and switching; The time difference is calculated based on the first timestamp and the second timestamp; User behavior data is obtained based on user actions, number of actions, and time differences. The touch pressure value, sliding acceleration, touch position and duration during the user's touch process are obtained as touch pressure data.
[0013] Preferably, the step of adjusting the weights of the memory enhancement network based on the feedback information to optimize generation efficiency includes: The feedback information, along with the corresponding text information, memory feature information, and target image, are used as training sample pairs and input into the long-term user profile memory unit of the memory enhancement network. The low-rank adapter weight file is then incrementally and lightweightly fine-tuned to obtain updated weight parameters. The training samples are input into the short-term session memory unit of the memory enhancement network to adjust the retrieval priority and storage period of historical feature vectors, thereby obtaining updated feature retrieval weights.
[0014] Secondly, embodiments of the present invention provide a generation efficiency optimization device based on intent recognition and memory, the device comprising: The judgment module is used to input the original audio data into the intent recognition and analysis model to obtain the text information, urgency level and content complexity level corresponding to the original audio data; The target generation mode module is used to determine the target generation mode based on the urgency level and the content complexity level. The memory feature module is used to obtain memory feature information that matches the text information through a preset memory enhancement network; The generation module is used to generate target generative content based on the text information, memory feature information and the target generation pattern; The information acquisition module is used to acquire user feedback information on the target generative content; The optimization module is used to adjust the weights of the memory enhancement network based on the feedback information to optimize the generation efficiency.
[0015] Thirdly, embodiments of the present invention provide an electronic device, including: at least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method of the first aspect described above.
[0016] Fourthly, embodiments of the present invention provide a storage medium storing computer program instructions, which, when executed by a processor, implement the method of the first aspect described above.
[0017] In summary, the beneficial effects of the present invention are as follows: The generation efficiency optimization method, apparatus, and device based on intent recognition and memory provided in this invention involve inputting raw audio data into a pre-trained intent recognition and analysis model to obtain text information, urgency level, and content complexity level corresponding to the raw audio data; determining a target generation mode based on the urgency level and content complexity level; acquiring memory feature information matching the text information through a preset memory enhancement network; generating target generative content based on the text information, memory feature information, and the target generation mode; obtaining user feedback information on the target generative content; and adjusting the weights of the memory enhancement network based on the feedback information to optimize generation efficiency. This invention achieves refined quantitative analysis of user-generated request intent through a pre-trained intent recognition and analysis model, laying a core foundation for accurate matching of generation strategies. It determines the target generation mode based on urgency and content complexity levels, realizing intent-driven hierarchical generation, significantly improving generation response efficiency while effectively reducing the ineffective consumption of computing resources. The matching memory feature information, combined with text information and the target generation mode, generates target content, integrating contextual information and personalized features into the generated results to better meet user needs. By obtaining user feedback on the generated content and adjusting the weights of the memory-enhancing network accordingly, it achieves continuous adaptive evolution of the memory-enhancing network, optimizing memory matching and generation logic, and further improving generation efficiency and content quality. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, and these are all within the protection scope of the present invention.
[0019] Figure 1 This is a schematic diagram of the overall process of the generation efficiency optimization method based on intent recognition and memory in an embodiment of the present invention; Figure 2 This is another flowchart illustrating the optimization of generation efficiency based on intent recognition and memory in an embodiment of the present invention; Figure 3 This is a flowchart illustrating the determination of content complexity level in the generation efficiency optimization based on intent recognition and memory in an embodiment of the present invention. Figure 4 This is a schematic diagram of the generation efficiency optimization device based on intent recognition and memory according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0020] The features and exemplary embodiments of various aspects of the present invention will now be described in detail. To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only configured to explain the present invention and are not configured to limit the present invention. For those skilled in the art, the present invention can be practiced without some of these specific details. The following description of the embodiments is merely intended to provide a better understanding of the present invention by illustrating examples of the invention.
[0021] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0022] Example 1 Please see Figure 1This invention provides a method for optimizing generation efficiency based on intent recognition and memory, the method comprising: S1: Input the raw audio data into the pre-trained intent recognition and analysis model to obtain the text information, urgency level and content complexity level corresponding to the raw audio data; Specifically, raw audio data refers to the raw PCM audio stream of user speech captured by the device's microphone, which is an unprocessed speech signal. The intent recognition and analysis model is a pre-trained model deployed locally on the terminal device, integrating speech signal processing, lightweight natural language understanding, and multimodal intent classification technologies. It adopts a signal processing-feature extraction-fusion judgment architecture, comprehensively utilizing speech signal processing, lightweight natural language understanding, and multimodal intent classification technologies to achieve low-latency, high-accuracy intent recognition under the limited computing power of the terminal device, obtaining text commands, urgency levels, and content complexity levels. The intent recognition and analysis model runs locally on the terminal device, performing real-time analysis of the user's input speech commands and quantitatively evaluating them from two dimensions: urgency level and content complexity level. Urgency assessment is achieved by extracting speech rate, volume, and the presence of preset urgent keywords in the speech signal, such as "hurry up," "immediately," or "draw another one." When a speech rate exceeding 150% of the user's average speech rate or the presence of urgent keywords is detected, it is determined to be a high urgency level. Content complexity assessment uses lightweight natural language processing (NLP) to analyze instruction text and evaluate the number of entities it describes, the complexity of relationships between entities, and the density of unconventional adjectives. For example, "a dog" is considered low complexity, while "a Shiba Inu wearing a pilot's hat and sunglasses, driving a red race car" is considered high complexity.
[0023] S2: Determine the target generation mode based on the urgency level and content complexity level; Specifically, based on the intent recognition results, the system dynamically switches between multiple generation modes to achieve a precise balance between computational power and quality. For example, it selects between three modes: Speed Mode, Balanced Mode, and Fine Mode. Speed Mode does not call any generation model; instead, it directly retrieves the most recently generated result or a matching PNG template library from the local cache, performs a rapid affine transformation, and outputs it immediately. Affine transformations, such as rotation and scaling, have a target response time of less than one second. For example, Speed Mode is triggered when the original audio data corresponds to a high urgency level and a low complexity level (e.g., printing another image). Balanced Mode runs a lightweight ControlNet model locally. ControlNet is a conditionally controlled generation model based on diffusion models (such as StableDiffusion). By introducing multimodal input conditions (such as edge detection maps, depth maps, pose keypoints, etc.), it significantly improves the controllability and detail restoration capabilities of image generation. Through the ControlNet model, using previously generated images or latent features as conditions, it quickly performs specified style transfer, color replacement, or local editing, achieving efficient fine-tuning. For example, the balanced mode is triggered when the original audio data corresponds to a medium level of complexity, or when a request requires partial adjustments (e.g., changing the previous one to blue). The refined mode sends complete text prompts, enhanced with contextual and historical preference information, to a large cloud model for full, high-quality generation. The refined mode is triggered when the original audio data corresponds to a high level of complexity, or when a request requires creativity (e.g., a completely new and complex scene description). When urgency is high, locally cached generation templates or historical generation records are prioritized; when urgency is low, a deep generation process is initiated.
[0024] S3: Obtain memory feature information that matches the text information through a preset memory enhancement network; Specifically, the pre-defined memory enhancement network is a two-level memory mechanism built on the terminal device, including short-term conversation memory and long-term user profile memory. Memory feature information refers to historical generation-related features that match the semantics of the current text instruction, including historical generation latent space feature vectors in short-term memory and user style preference low-rank adapter weight information in long-term memory. For example, the text information is first semantically parsed to detect whether it contains pronouns (such as "like this," "that," "just now"). If it does, the retrieval mechanism of short-term conversation memory is triggered, and the corresponding historical latent space feature vector is matched in the conversation stack through cosine similarity calculation. At the same time, regardless of whether pronouns exist, the loading mechanism of long-term user profile memory is triggered based on the user identity confirmation result, and the low-rank adapter weight file exclusive to that user is retrieved. Finally, the retrieved short-term memory features are fused with the loaded long-term memory features to obtain complete memory feature information. The low-rank adapter weight file refers to a lightweight model weight file that is independently maintained and stored for each user on the interactive smart hardware terminal, based on low-rank adapter (LoRA) technology. It is the core carrier for realizing user long-term generation preference memory and personalized content generation. By fusing the retrieved short-term memory features with the loaded long-term memory features, the subsequent generated results have contextual continuity and better meet the user's needs.
[0025] S4: Generate target-generated content based on the text information, memory feature information, and target generation pattern; Specifically, target-generated content refers to AIGC content ultimately generated based on the current user request, such as images / stickers output by AI sticker printers or interactive learning machines. Text information refers to the user's text request after speech recognition, serving as the core semantic basis for the generated content. The targeted application of different generation modes reduces the generation response time for simple requests to less than one second, while complex requests achieve high-quality generation results. The integration of memory feature information ensures that the generated content closely matches user needs, improving generation efficiency and content quality.
[0026] S5: Obtain user feedback on the target generative content; Specifically, feedback information refers to the analysis results of various data collected through device sensors that reflect user satisfaction with generated content. This data may include feedback type and feedback score, serving as the core basis for subsequent optimization of the memory network. Simultaneously with the display of target-generated content, multiple sensors, cameras, touchscreens, and other devices are used to collect real-time user behavior feedback data, such as gaze duration, operational behavior data, head posture characteristics, repetition request intervals, and touch pressure data. The collected raw data is then analyzed according to preset judgment rules to determine the user's feedback type (positive or negative feedback), and a feedback score is quantified based on data indicators. Finally, the feedback type and feedback score are integrated into complete user feedback information and stored. This provides a basis for subsequent weight adjustments in the memory enhancement network, enabling adaptive evolution.
[0027] S6: Based on the feedback information, adjust the weights of the memory enhancement network to optimize the generation efficiency.
[0028] Specifically, weight adjustment refers to incremental modifications to the retrieval priority and storage period of short-term session memory units and the LoRA weight files of long-term user profile memory units in the memory enhancement network. This transforms real user feedback into the basis for optimizing the memory network, allowing it to continuously accumulate user preferences and correct memory matching logic, thereby making the subsequent generation process more accurate and better suited to user needs.
[0029] In one embodiment, see Figure 2 S1 includes: S11: Preprocess the original audio data to obtain the target audio data; Specifically, in the preprocessing process, the original audio data is pre-emphasized and then framed, with a frame length of about 20-30ms and a frame shift of about 10ms. The original audio data after framing is then windowed (Hamming window) to obtain the target audio data. This process eliminates noise and solves problems such as signal distortion, meeting the requirements of subsequent speech recognition and feature extraction.
[0030] S12: Based on the audio recognition algorithm, perform text recognition on the target audio data to obtain text information; Specifically, the audio recognition algorithm refers to a lightweight automatic speech recognition (ASR) algorithm deployed locally on the terminal, adapted to the limited computing power of the terminal device; the text command refers to readable text information converted from standardized audio data, which is the core basis for subsequent natural language understanding, keyword matching, and complexity analysis. In practice, the target audio data can be first input into a lightweight ASR algorithm model pre-trained locally on the terminal. The model extracts audio features through an acoustic model, and then converts the acoustic features into a text sequence through a language model. At the same time, semantic correction and error correction are performed on the text sequence, such as correcting slips of the tongue and unclear pronunciations in speech into standard text. The output text information is fluent and accurate, and the entire recognition process is completed locally on the terminal without calling cloud resources.
[0031] S13: Perform natural language understanding on the text information to obtain a semantic feature vector, wherein the semantic feature vector is used to represent the number of entities, the structural relationship features between entities, and the number of unconventional adjectives; Specifically, natural language understanding refers to lightweight natural language understanding technology deployed on terminals, which focuses on extracting semantic features from text. Semantic feature vectors are numerical vectors used to quantify the semantic features of text instructions. They contain feature information in three dimensions: the number of entities, the structural relationship features between entities, and the number of unconventional adjectives. The number of entities refers to the number of core objects directly related to content generation in the text information, such as the number of animals, items, scenes, and people. For example, if the text information is "draw a dog," it's a single generated entity with a quantity of 1. If it's "draw a Shiba Inu driving a red race car," there are two entities: the Shiba Inu and the race car, resulting in a quantity of 2. The structural relationship features between entities refer to the complexity of grammatical modifications, logical nesting, and action relationships between entities in the text information. For example, "draw a cat and a dog" has a simple parallel structural relationship feature type, a syntactic tree depth of 1, and a branching factor of 1. "Draw a Shiba Inu driving a race car" has a single-layer action modification structural relationship feature type, a syntactic tree depth of 2, and a branching factor of 1. "Draw a Shiba Inu wearing a pilot's hat and sunglasses driving a red race car" has a multi-layer nested modification combined with action relationships structural relationship feature type, a syntactic tree depth of 4, and a branching factor of 3. The number of unconventional adjectives refers to the number of creative adjectives in text information that modify entities and are rare in the basic corpus of everyday conversations. It is a core quantitative feature representing the creativity and personalization of instructions. The basic corpus of everyday conversations is a pre-set corpus containing common vocabulary from daily conversations. Text instructions are input into a lightweight model on the terminal. The model first performs word segmentation, part-of-speech tagging, and syntactic analysis to identify and count the entities in the text. Then, dependency parsing is used to identify the modification and association relationships between entities, obtaining structural association features between entities. Simultaneously, adjectives are extracted from the text, and the number of unconventional adjectives is counted. Finally, the features of the three dimensions are quantified into numerical values and converted into vector form.
[0032] S14: According to the preset urgency rating standard, the target audio data is rated for urgency to obtain an urgency level, wherein the urgency level includes high urgency and low urgency. Specifically, the preset urgency rating standard refers to the pre-existing local urgency judgment rules that integrate three dimensions: speech rate, volume, and keywords. For example, if the speech rate exceeds the user's baseline value by 150%, or the volume exceeds the quiet state by 200%, or if there are urgent keywords, it is judged as high urgency. The urgency level refers to the quantitative judgment result of the urgency of the user's request, including two levels: high urgency and low urgency.
[0033] S15: Perform fuzzy reasoning on the semantic feature vector to obtain the content complexity level, wherein the content complexity level includes low complexity, medium complexity and high complexity.
[0034] Specifically, the content complexity level refers to the judgment result of the semantic complexity of the text instruction, including three levels: low, medium, and high. By combining the ambiguity of text semantics and the real-time status of the terminal device, the content complexity is accurately and dynamically determined, providing a basis for the selection of subsequent generation modes.
[0035] In one embodiment, S14 includes: S141: Based on the preset dual thresholds, perform endpoint detection on the target audio data to obtain valid audio segments; Specifically, the preset dual thresholds refer to the high and low energy thresholds set for endpoint detection, which can be dynamically adjusted according to the real-time status of the audio signal. Endpoint detection (VAD) is a technique for identifying the start and end positions of speech in an audio stream. Valid segments refer to audio fragments containing valid user speech segmented from standardized audio data, excluding silence and background noise. In endpoint detection, the short-time energy (STE) and short-time zero-crossing rate (ZCR) of each frame are calculated, and dynamic high and low energy thresholds are set to accurately segment valid segments and filter out silence and background noise.
[0036] S142: Perform syllable identification and statistics on the effective syllable segment to obtain the number of syllables in the effective syllable segment; Specifically, the syllable count refers to the number of syllables contained in the user's speech within the effective audio segment, and is a core parameter for calculating speech rate. In speech rate calculation, the audio within the effective audio segment can be processed using a lightweight recurrent neural network (the encoder part of an RNN-T) and a simplified phoneme recognition engine based on connection-time classification (CTC) to calculate the number of syllables or phonemes in the speech. The lightweight recurrent neural network (RNN) encoder refers to a lightweight RNN network adapted to the computing power of the terminal, retaining only the core feature extraction function; the simplified phoneme recognition engine based on connection-time classification (CTC) refers to a lightweight phoneme recognition model that removes complex post-processing modules, used for rapid recognition of phonemes or syllables. The valid audio segments are input to a lightweight recurrent neural network encoder on the terminal. The encoder extracts and encodes the audio features of the valid audio segments to obtain a high-dimensional audio feature vector. This feature vector is then input to a simplified phoneme recognition engine based on connection-time classification. The engine uses the CTC algorithm to achieve unaligned phoneme recognition, quickly identifying all phonemes within the valid audio segments. Based on the correspondence between phonemes and syllables, the identified phonemes are converted into syllables, and finally, the number of syllables within the valid audio segments is counted. This achieves fast local syllable estimation with low response latency, meeting the requirements of real-time interaction.
[0037] S143: Calculate the real-time speech rate based on the ratio of the number of syllables to the effective speech duration; Specifically, effective speech duration refers to the total duration of the effective speech segments segmented in step 1, measured in seconds; real-time speech rate refers to the user's actual speaking speed during this speech request, measured in syllables per second, used to characterize the speed at which the user speaks. The final real-time speech rate is obtained by the ratio of the number of syllables to the effective speech duration.
[0038] S144: Based on the preset speech rate emergency assessment rules, the real-time speech rate, and the pre-collected user-personalized baseline speech rate, the speech rate emergency characteristics are obtained. Specifically, the preset speech rate urgency assessment rule refers to the pre-existing local rules for determining the urgency of speech rate. For example, if the real-time speech rate exceeds 150% of the user's personalized baseline speech rate, it is considered to have a high urgency characteristic. The user's personalized baseline speech rate can be pre-collected through several initial interactive learning sessions. When determining urgency, the current real-time speech rate is compared with the user's personal baseline speech rate. For example, if it exceeds 150% of the baseline speech rate, it is judged as having a high urgency. A threshold can also be used for a comprehensive judgment.
[0039] S145: Calculate the short-time energy value of each frame of the target audio data according to the short-time energy calculation algorithm; Specifically, the short-time energy calculation algorithm can be implemented through a dedicated computing unit deployed locally on the terminal for calculating the short-time energy of audio frames. The short-time energy calculation engine is the core module for volume recognition. The short-time energy value refers to the sum of the squares of the amplitudes of all speech sampling points within an audio frame, representing the energy level of the audio frame and subsequently used to calculate the volume. All frame data of the standardized audio data are input into the short-time energy calculation engine on the terminal. The engine processes each frame of audio data, sequentially calculating the sum of the squares of the amplitudes of all speech sampling points within each frame to obtain the corresponding short-time energy value for each frame. The short-time energy values of all frames are stored in frame order to provide a numerical basis for subsequent volume calculations.
[0040] S146: Perform weighted filtering and logarithmic transformation on the short-time energy value to obtain the real-time volume; Specifically, in volume recognition, a short-time energy calculation engine is used to directly calculate the short-time energy of each frame of the segmented speech signal. After weighted filtering and logarithmic transformation, the energy value yields a volume estimate in decibels. High-urgency requests are usually accompanied by a significant increase in voice volume, such as exceeding 200% of the average volume in a quiet state.
[0041] S147: Based on the real-time volume and the preset volume evaluation rules, obtain the volume emergency characteristics; Specifically, the preset volume evaluation rules refer to the volume urgency determination rules that are pre-existing locally. For example, "if the real-time volume exceeds 200% of the average volume when the user is quiet, it is determined to be a high volume urgency feature." The volume urgency feature refers to the determination result of whether the real-time volume represents an urgent request, which is one of the bases for determining the urgency level.
[0042] S148: Obtain the keyword emergency features based on the preset emergency keyword library and the text information; Specifically, the pre-defined emergency keyword library is a set of high-urgency trigger words predefined for interactive scenarios, such as "hurry up," "immediately," "again," "also," and their common variations, used as a vocabulary for matching. For identifying the urgent features of keywords, a lightweight word graph matching algorithm is used, specifically through the pre-defined emergency keyword library, real-time matching of text information, and fusion triggering. In real-time matching, after the system automatically recognizes the user's speech (ASR) and converts it into text commands, it does not wait for the complete sentence to end, but performs sliding window matching on the text stream. It utilizes a Tritree data structure to achieve fast matching of multi-pattern strings, quickly detecting whether any emergency keywords appear in the text stream. The fusion triggering mechanism means that when at least one emergency keyword is detected, a keyword emergency signal is generated as a comprehensive decision on the urgency of the speech.
[0043] S149: Based on the urgent characteristics of speech rate, urgent characteristics of volume, and preset urgent rating criteria, the urgent level is obtained.
[0044] Specifically, for example, a preset urgency rating standard is: if there is an urgent speech rate or an urgent volume, it is judged as high urgency; if neither an urgent speech rate nor an urgent volume occurs, it is judged as high urgency.
[0045] In one embodiment, see Figure 3 S15 includes: S151: Normalize the number of entities, the structural association features between entities, and the number of unconventional adjectives in the semantic feature vector to obtain each confidence score; Specifically, semantic feature vectors of text information are obtained after lightweight natural language processing (NLP) analysis. Further judgment employs a three-level hybrid technique: normalized quantization, fuzzy logic reasoning, and reinforcement learning boundary adjustment, rather than simply relying on thresholds. The semantic feature vectors output by the NLP module are first converted into continuous confidence scores between 0 and 1 through a normalized quantization function. This includes entity count quantization, relational complexity quantization, and creativity density quantization. In entity count quantization, a saturation function-based normalization is primarily used, setting a saturation value specific to the application scenario. For example, setting 5 entities as the saturation value means that when the number of entities is 1, the score is 0; when the number of entities is 3, the score is 0.5; and when the number of entities is ≥5, the score saturates to 1.0. This avoids the excessive mathematical inflation of simple requests (2 entities) and complex requests (10 entities). For quantifying relation complexity, a weighted calculation based on the depth and branching factor of the dependency syntax tree is primarily used, rather than simply counting the number of relations. The dependency syntax tree is recursively traversed; if the sentence is a simple parallel structure with a shallow tree depth, the score is low; if the sentence is a multi-level nested decoration with a deep tree depth (>3), the score is high. Finally, the sigmoid function is used to map the depth and branches to the 0-1 range to simulate the non-linear complexity of relations. For quantifying creativity density, a density calculation based on the inverse document frequency (IDF) set is used, which can be matched and quantified according to a pre-set "basic corpus" (containing common vocabulary from daily conversations). If an adjective (such as "red") is common in the basic corpus, its IDF value is low and its weight is low; if an adjective (such as "cyberpunk" or "liquid metal") is rare in the basic corpus, its IDF value is high and its weight is high. Finally, the output range is controlled by the tanh function, so that sentences containing multiple rare creative words can get significantly higher scores.
[0046] S152: Based on the preset membership function and the preset complexity scoring rule library, perform fuzzy reasoning and defuzzification on each confidence score to obtain a continuous complexity score. Specifically, in the fuzzy logic reasoning layer, human perception of complexity is fuzzy. A request with many entities but extremely simple relationships, such as listing cats, dogs, birds, and fish, is not complex; a request with only one entity but extremely complex modifying relationships, such as a cat scratching its left earlobe with its hind paw, is very complex. Threshold logic struggles to handle these edge cases. The first step is to input three consecutive scores into the fuzzy logic system and define a membership function, defining a fuzzy set (high, medium, low) for each input. Then, a rule base is embedded. For example: if IF S_ent is low, AND S_rel is low, AND S_cre is low, then the overall complexity is low; if IF S_rel is high, OR S_cre is high, then the overall complexity is high, even if there is only one entity, as long as the modification is extremely complex, it is judged as high; if IF S_ent is medium, AND S_rel is low, AND S_cre is low, then the overall complexity is medium; if IF S_ent is high, AND S_rel is low, AND S_cre is low, then the overall complexity is medium; if IF S_ent is high, AND S_rel is low, AND S_cre is low, then the overall complexity is medium. S_cre is low, THEN is medium (simply listing items is not considered high, only medium); finally, using the centroid method, based on the fuzzy results obtained from fuzzy reasoning, a continuous complexity score of 0~100 is calculated for defuzzification.
[0047] S153: Based on the continuous complexity score and the preset dynamic complexity decision threshold, the content complexity level is obtained.
[0048] Specifically, the dynamic complexity decision threshold refers to a threshold dynamically adjusted based on the real-time status of the device, used to determine low / medium / high complexity. In practical applications, the current available computing power and network latency of edge devices are monitored. Computing power data includes CPU utilization and remaining memory; network latency information includes network packet loss rate and transmission latency. When network conditions are excellent and battery power is sufficient, the decision boundary shifts to the right, meaning a higher complexity score is required to invoke the cloud, encouraging local balancing mode to handle more medium-sized tasks. When network congestion and low battery power occur, the decision boundary shifts to the left, meaning a lower complexity score may trigger the cloud, as local computing power needs to be conserved. Matching continuous complexity scores with the dynamic decision boundary yields the content complexity level, providing a complexity basis for subsequent generation mode selection. The corresponding generation mode is matched based on the complexity level.
[0049] In one embodiment, S2 includes: S21: When the content complexity level is low complexity and the urgency level is high urgency, the target generation mode is determined to be the rapid generation mode, wherein the rapid generation mode refers to the generation mode that directly calls the local template library without calling any model. Specifically, the ultra-fast generation mode refers to directly calling the local template library, performing geometric transformations, and outputting the result with a response time of less than 1 second. For example, when the speech rate is detected to be 150% higher than the baseline and the ASR recognition result contains keywords such as "again" or "previous image," the content complexity level is low and the urgency level is high. The ultra-fast mode is then used to retrieve the most recently generated final image or matching template from the local cache, call a lightweight image processing library to perform random, small-amplitude rotations and scaling transformations, and output a new sticker within 1 second, perfectly satisfying the child's immediate psychological need. Low-complexity requests have simple semantics (such as "draw another one" or "print it quickly"), requiring no complex creative generation or detailed rendering; the terminal can complete the processing locally. High urgency, on the other hand, represents a strong need for immediate feedback from the user (such as a child urgently demanding the result), where the requirement for response speed is far higher than the refinement of the generated content.
[0050] S22: When the content complexity level is low complexity and the urgency level is low urgency, the target generation mode is determined to be a balanced generation mode, wherein the balanced generation mode refers to the generation mode that calls the local lightweight model. Specifically, the balanced generation mode runs a lightweight model locally, such as the ControlNet model, using previously generated images or latent features as conditions to quickly perform specified style transfers, color replacements, or local editing, achieving efficient fine-tuning. While low-complexity requests do not require high-computing power generation, low urgency means users do not have an urgent need for response speed. If the ultra-fast generation mode is still used to perform only simple template transformations, the generated content may be monotonous and lack personalization. By using a local lightweight model, basic personalized adjustments can be achieved while ensuring that the generated content accurately matches simple semantics.
[0051] S23: When the content complexity level is medium complexity and the urgency level is low urgency, the target generation mode is determined to be a balanced generation mode. Specifically, medium-complexity requests require basic local editing, feature modification, or style adjustment (such as "adding a red collar to the kitten"), but do not reach the level of needing to generate a full large model in the cloud; low-urgency requests have ample time to complete processing locally on the terminal, without sacrificing quality for speed, making a balanced generation mode more suitable.
[0052] S24: When the content complexity level is medium complexity and the urgency level is high urgency, the target generation mode is determined to be the rapid generation mode. Specifically, while medium complexity requires modifications and adjustments, high urgency indicates that the user's core demand is to get results quickly, rather than extreme detail and precision. If the balanced generation mode is selected, the lightweight model on the terminal's local machine still takes some time to process medium complexity, which cannot meet the immediate needs. On the other hand, the ultra-fast generation mode can call the matching medium complexity template or historical generated content in the local cache, and output it quickly after minor lightweight adjustments, prioritizing the user's urgent interaction needs, especially the behavioral characteristics of users such as children who need immediate gratification.
[0053] S25: When the content complexity level is high complexity, the target generation mode is determined to be a fine generation mode, wherein the fine generation mode refers to the generation mode using a large cloud model; Specifically, the refined generation mode refers to sending complete text prompts, enhanced with contextual and historical preference information, to a cloud-based large-scale model for full-scale, high-quality generation to meet users' creative needs. The cloud-based large-scale model possesses ample computing power and comprehensive generation capabilities, enabling full-scale, high-quality generation of highly complex content, ensuring the creativity of the generated results and satisfying users' advanced generation requirements.
[0054] S26: When the content complexity level is high complexity and the urgency level is high urgency, the target generation mode is determined to be the balanced generation mode.
[0055] Specifically, if the fine-grained generation mode is invoked according to the high-complexity rule, the computation and network transmission of the large cloud model will cause significant latency, failing to meet the user's urgent needs. If the high-urgency rule is invoked according to the rapid generation mode, simple template transformations will be completely unable to match the high-complexity semantics, resulting in generated content that deviates significantly from the user's requirements. The balanced generation mode, based on the terminal's local lightweight model, invokes locally cached similar high-complexity features and historical latent space features to complete the lightweight and rapid generation of high-complexity requests. While ensuring that the generated content basically conforms to the high-complexity semantics, it minimizes response latency, satisfying the user's urgent acquisition needs while ensuring that the generated results basically match their creative requirements, achieving a compromise between speed and quality. The three generation modes have preset computing power budget units. For example, the rapid mode budget is <0.5 units, the balanced mode budget is 1.0 unit, and the fine-grained mode budget is 3.0 units. The decision engine prioritizes the path with the lower budget while satisfying the intent.
[0056] In one embodiment, S5 includes: S51: Obtain user behavior feedback data on target generated content, wherein the behavior feedback data includes gaze duration, operation behavior data, head posture characteristics, repetitive request interval and touch pressure data; Specifically, behavioral feedback data refers to the raw data collected by various sensors on the terminal device that reflects the user's real experience with the generated content, and is the basis for subsequent analysis of feedback information; gaze duration refers to the cumulative time the user's head is facing the generated content; head posture characteristics refer to data such as the user's head yaw angle, pitch angle, and shoulder key point coordinates collected by the RGB camera; touch pressure data refers to the pressure value of the user touching the screen collected by the touch screen pressure sensor.
[0057] S52: Based on preset judgment rules and behavioral feedback data, obtain the user's feedback type and feedback score, wherein the feedback type includes positive feedback and negative feedback; Specifically, the preset judgment rules refer to the feedback judgment rules formulated for various behavioral feedback data, including the judgment conditions for positive and negative feedback and the quantitative standards for feedback scores; feedback type refers to the qualitative judgment of user satisfaction, divided into positive feedback (satisfaction) and negative feedback (dissatisfaction); feedback score refers to the quantitative quantification of user satisfaction, which can be a value between 0 and 100, with higher scores indicating higher satisfaction. For example, when the head orientation angle is less than a certain threshold, such as yaw ±15° and pitch ±10°, it is judged as gaze, and then a timer is started to accumulate the duration of consecutive frames that meet the conditions. When the accumulated duration exceeds 3 seconds, a long-term positive feedback signal is output. Facial expressions are judged by analyzing the changes in head posture sequences. Each frame outputs the head's yaw angle, pitch angle, and shoulder key points. When the absolute value of the yaw angle rapidly increases from <15° to >45°, accompanied by a change in the body's torso orientation, a head-turning event is considered detected. When the distance between the head and shoulders shortens by >20% and the pitch angle drops, a backward tilting time is considered detected. If the above behaviors occur within 3 seconds after the image is displayed and last for >0.5 seconds, they are judged as "negative avoidance behavior," equivalent to negative facial expression feedback.
[0058] S53: Based on the feedback score and feedback type, obtain the user's feedback information on the target generative content.
[0059] Specifically, by comprehensively considering feedback scores and feedback types, user feedback information on the target generative content is obtained.
[0060] In one embodiment, S51 includes: S511: Acquire several consecutive frames of images from the user; Specifically, the image acquisition process is triggered the instant the target-generated content is rendered and displayed on the terminal device screen. For example, the terminal's built-in ordinary RGB camera is used to continuously acquire an image stream at a fixed frame rate (e.g., 30 frames / second). After acquiring several consecutive frames of images through the ordinary RGB camera, head pose estimation and continuous frame tracking analysis are then performed.
[0061] S512: Input several of the images into a preset human posture detection model to obtain the head yaw angle, head pitch angle and coordinates of key human points; Specifically, the human pose detection model refers to an ultra-lightweight pose recognition model deployed on the terminal, such as the MediaPipe Pose model (a high-precision human tracking model). This model can convert images into quantifiable pose values, enabling rapid detection of key human points. Head yaw angle refers to the angle of left-right head rotation, with the device screen as the front; leftward yaw is negative, rightward yaw is positive, and the range is -90° to 90°. Head pitch angle refers to the angle of up-down head rotation, with the horizontal as the reference point; downward pitch is negative, upward pitch is positive, and the range is -90° to 90°. Human key point coordinates refer to outputting only the core key points related to the head and shoulders, which can include the two-dimensional pixel coordinates of the nose center, the left and right eye centers, and the left and right shoulder centers, for example, the nose center coordinates (120, 80). The model does not perform face recognition, extract facial feature vectors, or build a facial biometric database. All processing is completed locally on the terminal device; image processing is real-time, without storage or transmission. The model performs feature extraction, key point detection, and angle calculation for each frame of image. First, it identifies the head and shoulder key points in the image and outputs their pixel coordinates. Then, based on the positional relationship of the key points, it calculates the head yaw angle and head pitch angle corresponding to each frame. The angle data of all frames and the key point coordinate data are associated and stored according to the timestamp to ensure real-time synchronization with image acquisition.
[0062] S513: Based on the difference in head yaw angle and head pitch angle between adjacent frames and a preset angle difference threshold, determine whether the user is looking at the generated content. Specifically, the angle difference between adjacent frames refers to the difference in head yaw angle and head pitch angle between two consecutive frames, reflecting the magnitude of head angle change. The preset angle difference threshold is a critical value set to determine whether the head is stable. For example, when the yaw angle is ±15° and the pitch angle is ±10° for multiple consecutive frames, it is determined that the user is looking at the generated content. This achieves effective determination of gaze state without collecting facial features.
[0063] S514: When it is determined that the user is gazing at the generative content, obtain the gazing duration; Specifically, if it is determined that the user is gazing at the generated content, the duration of consecutive frames that meet the cooking condition is accumulated according to a timer to obtain the gazing duration and quantify the user's attention to the generated content.
[0064] S515: Obtain head posture features based on the coordinates of the human body key points and the preset posture judgment template; Specifically, the posture judgment template refers to a set of head posture judgment rules preset in the terminal based on the correlation between human posture features and user feedback. It includes judgment conditions and feature definitions for different head posture events. Head posture features refer to user head posture event information extracted by analyzing changes in the coordinates of key human points, including posture event type, trigger time, and duration, reflecting the user's implicit emotional feedback to the generated content. Subsequently, for negative expression judgment, a continuous frame image stream is read using a regular RGB camera, and expression judgment is made by analyzing the changes in the head posture sequence. Each frame outputs the head's yaw angle, pitch angle, and shoulder key points. When the absolute value of the yaw angle rapidly increases from <15° to >45°, accompanied by a change in the body's torso orientation, a head-turning event is detected. When the distance between the head and shoulders shortens by >20%, and the pitch angle decreases downwards, a backward tilting event is detected. If these behaviors occur within 3 seconds of image display and last for >0.5 seconds, they are judged as negative avoidance behavior, equivalent to negative expression feedback.
[0065] S516: Obtain the first timestamp of the target generated content display and the second timestamp of the user operation, wherein the user operation includes skipping, saving and switching; Specifically, the first timestamp refers to the moment when the target generated content is fully displayed on the terminal screen, the time point recorded by the system; the second timestamp refers to the moment when the user triggers operations such as skip, save, or switch, the time point recorded by the system; the skip operation refers to the operation of the user clicking buttons such as "skip," "next," or "back" or swiping the screen to quickly skip the current generated content; the save operation refers to the operation of the user clicking buttons such as "save," "print," or "favorite" to save the current generated content to the device or output; the switch operation refers to the operation of the user clicking buttons such as "change," "regenerate," or "adjust" to request the generation of new content to replace the current content.
[0066] S517: Calculate the time difference based on the first timestamp and the second timestamp; Specifically, the time difference is obtained by calculating the difference between the second timestamp and the first timestamp.
[0067] S518: Obtain operation behavior data based on user actions, number of actions, and time difference; Specifically, operational behavior data refers to structured data formed by integrating user operation types, number of operations, and operation time differences. This data quantifies users' explicit feedback to generated content and includes basic operational information and feedback tendency judgment information. For example, the time difference is obtained by acquiring button click time, page switching time, and timestamps. The first timestamp of image generation and display, and the second timestamp of the user's triggering of skip-related operation words are also acquired, and the time difference is calculated. If the time difference is less than 1 second, it is judged as a rapid skip, and a negative feedback signal is output.
[0068] S519: Obtain the touch pressure value, sliding acceleration, touch position and duration during the user's touch process as touch pressure data.
[0069] Specifically, touch pressure value refers to the amount of pressure applied to the screen when a user touches the terminal's touchscreen, including peak and average pressure values, measured in Newtons (N); swipe acceleration refers to the magnitude of acceleration when a user performs a swipe operation on the touchscreen, reflecting the speed change of the swipe operation, measured in m / s²; touch position refers to the two-dimensional pixel coordinates of the touch point on the screen when the user touches the touchscreen; touch duration refers to the length of time from the start to the end of a single touch or swipe operation; touch pressure data refers to structured data formed by integrating touch pressure value, swipe acceleration, touch position, and touch duration, which is an important supplementary data reflecting the intensity of the user's emotions and intentions during operation. Touch pressure data is acquired through a touchscreen / button time listener to determine whether there is any rapid skipping. In determining user feedback to the generated results, the evaluation criteria also include the interval between repeated requests. If a user sends a highly semantically similar request immediately after generating an image, it is considered strong positive feedback. Physical operation force and speed are also considered. For devices with touch screens, indicators such as touch pressure or sliding acceleration are collected as evaluation criteria. The evaluation is based on speech language features. Emotional speech recognition is performed on the ambient speech within 5 seconds after the generated result is displayed, and unconscious speech made by the user after seeing the generated result is collected.
[0070] In one embodiment, S6 includes: S61: The feedback information, along with the corresponding text information, memory feature information, and target image, are used as training sample pairs and input into the long-term user profile memory unit of the memory enhancement network. The low-rank adapter weight file is then incrementally and lightweightly fine-tuned to obtain updated weight parameters. Specifically, locally on the terminal, the basic generative model is continuously and lightweightly fine-tuned using a low-rank adapter (LoRA). The low-rank adapter weight file continuously learns and records the user's long-term, stable style preferences, such as a preference for blue tones, frequent use of dinosaur themes, or a liking for cartoon styles. Training sample pairs are input into the long-term user profile memory unit, which uses the low-rank adapter weight file as its core. Based on feedback scores, the fine-tuning magnitude is determined, and small modifications are made to the low-rank matrices corresponding to the style / entity features relevant to this generation: positive feedback increases the weight of the corresponding feature, while negative feedback decreases the weight of the corresponding feature. The entire fine-tuning process is an incremental, lightweight operation, involving only the low-rank adapter weight file and not the full training of the basic generative model. After fine-tuning, updated low-rank adapter weight parameters are obtained, serving as the basis for the user's unique long-term style preferences. During each generation, the user's personal low-rank adapter weights are automatically loaded, ensuring that the output naturally aligns with their historical preferences. For example, during long-term use, the system detects that users repeatedly generate images related to the starry sky. When such images are generated in the cloud, the system uses the current prompt and the generated result as training pairs, performing a small-scale low-rank adapter weight adjustment locally, and updating the weights of the user's personal low-rank adapter weight file. Months later, if a user simply enters "beautiful night," and the generation model with their personal low-rank adapter weights is loaded, it will also tend to automatically add starry sky elements, achieving personalization.
[0071] S62: Input the training samples into the short-term session memory unit of the memory enhancement network, adjust the retrieval priority and storage period of historical feature vectors, and obtain updated feature retrieval weights.
[0072] Specifically, historical feature vectors refer to the latent space feature vectors generated in the previous sessions, stored in the session stack, and are the core basis for achieving contextual association; retrieval priority refers to the degree to which historical feature vectors are retrieved during semantic matching; storage period refers to the retention time of historical feature vectors in the session stack; feature retrieval weight refers to the weight parameter used to adjust the retrieval priority of historical feature vectors, with higher weights resulting in higher retrieval priority. Short-term session memory refers to maintaining a session stack in device memory, dynamically storing the latent space feature vectors of images generated within the current session period, and supporting referential resolution and contextual association. When referential resolution words (e.g., "like this," "that," "just now") are detected, the corresponding historical features are automatically retrieved from the stack and injected as generation conditions, ensuring a high degree of consistency in the context. First, the feedback type and score in the training sample pairs are analyzed, along with the corresponding historical feature vectors (short-term memory features used in this generation). Based on the feedback results, the retrieval priority and storage period of these historical feature vectors are adjusted: if the feedback is positive and the score is high, the retrieval priority or feature retrieval weight of the feature vector is increased, and its storage period in the session stack is extended; if the feedback is negative, the retrieval priority or feature retrieval weight of the feature vector is decreased, its storage period is shortened, or it may even be removed directly from the session stack. Simultaneously, based on the feedback of all historical feature vectors, the feature retrieval weights of all feature vectors in the session stack are recalculated and updated, forming a new retrieval priority ranking to provide a basis for subsequent memory retrieval. For example, if the user's previous dialogue was: "Generate an astronaut cat," the latent feature vector generated this time is stored in the short-term memory stack. Then, if the user says: "Add a red flag to it," the NLP module recognizes "it" as a referent and immediately retrieves the latent feature vector from the memory stack. The system enters balanced mode, inputting the latent feature vector and the new text information "add a red flag" into the local ControlNet model. Based on the latent feature vector, the model quickly generates an image of an astronaut cat with the red flag added. The entire process is completed locally, taking about 1.5 seconds, and the style is completely consistent.
[0073] Example 2 Please see Figure 4 This invention provides a generation efficiency optimization device based on intent recognition and memory, the device comprising: The judgment module is used to input the original audio data into the intent recognition and analysis model to obtain the text information, urgency level and content complexity level corresponding to the original audio data; The target generation mode module is used to determine the target generation mode based on the urgency level and the content complexity level. The memory feature module is used to obtain memory feature information that matches the text information through a preset memory enhancement network; The generation module is used to generate target generative content based on the text information, memory feature information and the target generation pattern; The information acquisition module is used to acquire user feedback information on the target generative content; The optimization module is used to adjust the weights of the memory enhancement network based on the feedback information to optimize the generation efficiency.
[0074] It should be noted that each module and unit in the generation efficiency optimization device based on intent recognition and memory in this embodiment corresponds one-to-one with each step in the generation efficiency optimization method based on intent recognition and memory in the aforementioned embodiment. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned generation efficiency optimization method based on intent recognition and memory, and will not be repeated here.
[0075] Example 3 Furthermore, the generation efficiency optimization method based on intent recognition and memory in this embodiment of the invention can be implemented by an electronic device. Figure 5 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention is shown.
[0076] Electronic devices may include processors and memory storing computer program instructions.
[0077] Specifically, the processor may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement embodiments of the present invention.
[0078] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0079] Computer-readable media include both permanent and non-permanent, removable and non-removable media, which can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient media, such as modulated communication signals and carrier waves.
[0080] The processor reads and executes computer program instructions stored in memory to implement any of the generation efficiency optimization methods based on intent recognition and memory in the above embodiments.
[0081] In one example, the electronic device may also include a communication interface and a bus. For example, Figure 5 As shown, the processor 401, memory 402, and communication interface 403 are connected through bus 410 and complete communication with each other.
[0082] The communication interface is mainly used to enable communication between various modules, devices, units and / or equipment in the embodiments of the present invention.
[0083] A bus, including hardware, software, or both, couples components of an electronic device together. For example, and not limitingly, a bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, a bus may include one or more buses. While specific buses are described and illustrated in embodiments of the invention, the invention contemplates any suitable bus or interconnect.
[0084] Example 4 Furthermore, in conjunction with the generation efficiency optimization method based on intent recognition and memory in the above embodiments, this invention can be implemented using a computer-readable storage medium. This computer-readable storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the generation efficiency optimization methods based on intent recognition and memory in the above embodiments.
[0085] In summary, the generation efficiency optimization method, apparatus, and device based on intent recognition and memory provided in this invention achieve refined quantitative analysis of user-generated request intent through a pre-trained intent recognition and analysis model, laying a core foundation for accurate matching of generation strategies. By determining the target generation mode based on urgency and content complexity levels, it achieves intent-driven hierarchical generation, significantly improving generation response efficiency while effectively reducing the ineffective consumption of computing resources. The matching memory feature information, combined with text information and the target generation mode, generates target content, integrating contextual information and personalized features into the generated results, better meeting user needs. By obtaining user feedback on the generated content and adjusting the weights of the memory-enhancing network accordingly, it achieves continuous adaptive evolution of the memory-enhancing network, optimizing memory matching and generation logic, further improving generation efficiency and content quality.
[0086] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.
[0087] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0088] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0089] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0090] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0091] It should also be noted that the exemplary embodiments mentioned in this invention describe methods or systems based on a series of steps or apparatus. However, this invention is not limited to the order of the steps described above; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0092] The above description is merely a specific embodiment of the present invention. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the protection scope of the present invention.
Claims
1. A method for optimizing generation efficiency based on intent recognition and memory, characterized in that, The method includes: The raw audio data is input into a pre-trained intent recognition and analysis model to obtain text information, urgency level, and content complexity level corresponding to the raw audio data. Based on the urgency level and content complexity level, determine the target generation mode; Memory feature information matching the text information is obtained through a preset memory enhancement network; Based on the text information, memory feature information, and target generation pattern, generate target-generated content; Obtain user feedback on the target generative content; Based on the feedback information, the weights of the memory enhancement network are adjusted to optimize the generation efficiency.
2. The method according to claim 1, characterized in that, The process of inputting the raw audio data into a pre-trained intent recognition and analysis model yields text information, urgency level, and content complexity level corresponding to the raw audio data, including: The original audio data is preprocessed to obtain the target audio data; Based on the audio recognition algorithm, the target audio data is subjected to text recognition to obtain text information; Natural language understanding is performed on the text information to obtain a semantic feature vector, wherein the semantic feature vector is used to represent the number of entities, the structural relationship features between entities, and the number of unconventional adjectives; According to the preset urgency rating criteria, the target audio data is rated for urgency to obtain an urgency level, wherein the urgency level includes high urgency and low urgency. Fuzzy reasoning is performed on the semantic feature vector to obtain the content complexity level, wherein the content complexity level includes low complexity, medium complexity and high complexity.
3. The method according to claim 2, characterized in that, The step of rating the target audio data according to a preset urgency rating standard to obtain an urgency level includes: Based on preset dual thresholds, endpoint detection is performed on the target audio data to obtain valid audio segments; Syllable identification and counting are performed on the effective syllable segments to obtain the number of syllables within the effective syllable segments; The real-time speech rate is calculated based on the ratio of the number of syllables to the effective speech duration. Based on the preset speech rate emergency assessment rules, the real-time speech rate, and the pre-collected user-personalized baseline speech rate, the speech rate emergency characteristics are obtained. The short-time energy value of each frame of the target audio data is calculated based on the short-time energy calculation algorithm. The short-time energy value is weighted, filtered, and logarithmically transformed to obtain the real-time volume. Based on the real-time volume and the preset volume evaluation rules, the volume emergency characteristics are obtained; Based on the preset emergency keyword database and the text information, the emergency characteristics of the keywords are obtained; An urgency level is obtained based on the urgency characteristics of speech rate and volume, as well as the preset urgency rating criteria.
4. The method according to claim 2, characterized in that, The process of performing fuzzy reasoning on the semantic feature vector to obtain the content complexity level includes: The number of entities, the structural association features between entities, and the number of unconventional adjectives in the semantic feature vector are normalized respectively to obtain each confidence score; Based on the preset membership function and the preset complexity scoring rule library, fuzzy reasoning and defuzzification are performed on each confidence score to obtain a continuous complexity score. The content complexity level is obtained based on the continuous complexity score and the preset dynamic complexity decision threshold.
5. The method according to claim 2, characterized in that, The step of determining the target generation mode based on the urgency level and content complexity level includes: When the content complexity level is low and the urgency level is high, the target generation mode is determined to be the rapid generation mode. The rapid generation mode refers to the generation mode that directly calls the local template library without calling any model. When the content complexity level is low and the urgency level is low, the target generation mode is determined to be a balanced generation mode, wherein the balanced generation mode refers to the generation mode that calls the local lightweight model. When the content complexity level is medium complexity and the urgency level is low urgency, the target generation mode is determined to be a balanced generation mode. When the content complexity level is medium complexity and the urgency level is high urgency, the target generation mode is determined to be the rapid generation mode. When the content complexity level is high complexity, the target generation mode is determined to be a refined generation mode, wherein the refined generation mode refers to the generation mode using a large cloud model; When the content complexity level is high complexity and the urgency level is high urgency, the target generation mode is determined to be the balanced generation mode.
6. The method according to claim 1, characterized in that, The step of obtaining user feedback information on the target generative content includes: Acquire user behavior feedback data on target generated content, wherein the behavior feedback data includes gaze duration, operation behavior data, head posture features, repetition request interval and touch pressure data; Based on preset judgment rules and behavioral feedback data, the user's feedback type and feedback score are obtained, wherein the feedback type includes positive feedback and negative feedback; Based on the feedback score and feedback type, the user's feedback information on the target generative content is obtained.
7. The method according to claim 6, characterized in that, The acquisition of user behavioral feedback data on the target generated content includes: Acquire several consecutive frames of images from the user; Input several of the images into a preset human posture detection model to obtain the head yaw angle, head pitch angle and coordinates of key human points; Based on the difference in head yaw angle and head pitch angle between adjacent frames and a preset angle difference threshold, it is determined whether the user is looking at the generated content. When it is determined that a user is gazing at the generative content, the gazing duration is obtained; Based on the coordinates of the key human body points and the preset posture judgment template, the head posture features are obtained; Obtain a first timestamp of the target generated content display and a second timestamp of the user operation, wherein the user operation includes skipping, saving, and switching; The time difference is calculated based on the first timestamp and the second timestamp; User behavior data is obtained based on user actions, number of actions, and time differences. The touch pressure value, sliding acceleration, touch position and duration during the user's touch process are obtained as touch pressure data.
8. The method according to claim 1, characterized in that, The step of adjusting the weights of the memory enhancement network based on the feedback information to optimize generation efficiency includes: The feedback information, along with the corresponding text information, memory feature information, and target image, are used as training sample pairs and input into the long-term user profile memory unit of the memory enhancement network. The low-rank adapter weight file is then incrementally and lightweightly fine-tuned to obtain updated weight parameters. The training samples are input into the short-term session memory unit of the memory enhancement network to adjust the retrieval priority and storage period of historical feature vectors, thereby obtaining updated feature retrieval weights.
9. A generation efficiency optimization device based on intent recognition and memory, characterized in that, The device includes: The judgment module is used to input the original audio data into the intent recognition and analysis model to obtain the text information, urgency level and content complexity level corresponding to the original audio data; The target generation mode module is used to determine the target generation mode based on the urgency level and the content complexity level. The memory feature module is used to obtain memory feature information that matches the text information through a preset memory enhancement network; The generation module is used to generate target generative content based on the text information, memory feature information and the target generation pattern; The information acquisition module is used to acquire user feedback information on the target generative content; The optimization module is used to adjust the weights of the memory enhancement network based on the feedback information to optimize the generation efficiency.
10. An electronic device, characterized in that, include: At least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method as described in any one of claims 1-7.