An object representation generation method, system and computer readable storage medium

By employing a multi-level priority strategy and a feedback learning mechanism, the intelligent interaction system addresses the failure issue caused by a single data source when acquiring object representations. It achieves multimodal output and adaptive optimization, thereby enhancing user experience and interaction effects.

CN122310418APending Publication Date: 2026-06-30亓泽辰
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
亓泽辰
Filing Date
2026-03-31
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing intelligent interaction systems rely on a single data source when processing object representations, leading to acquisition failures, lack of learning mechanisms, rigid resource scheduling, inability to meet multimodal interaction needs, and poor user experience.

Method used

It employs a multi-level priority strategy to acquire object representations, including local matching, online acquisition, and content generation. Combined with a feedback learning mechanism, it dynamically adjusts the acquisition strategy, supports multimodal output such as visual, audio, and tactile, and applies preset style parameters to ensure consistency.

Benefits of technology

It improves the success rate and coverage of object representation acquisition, optimizes resource scheduling, enhances user experience, and enables multimodal interaction and adaptive optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122310418A_ABST
    Figure CN122310418A_ABST
Patent Text Reader

Abstract

With the development of artificial intelligence and interaction technologies, systems need to generate representations of various objects based on user input or context. In various interactive systems, users often expect the system to generate representations related to specific objects based on dialogue or instructions. However, existing technologies suffer from drawbacks such as a single acquisition method, lack of feedback learning, single representation modality, lack of flexible scheduling, and lack of linkage with actions. Therefore, a method is needed that can flexibly adopt multi-level strategies to acquire object representations, learn and optimize from user feedback, and support multi-modal representation generation to improve the intelligence of interaction and user experience. This application provides an object representation generation method, system, and computer-readable storage medium, aiming to ensure the success rate of acquisition through a multi-level priority strategy, dynamically optimize the acquisition strategy based on user feedback, and support multiple representation modalities such as vision, audio, touch, and action to achieve an efficient, personalized, and multi-modal interactive experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to artificial intelligence, human-computer interaction, multimodal information processing, and content generation. Specifically, it relates to an object representation generation method, system, and computer-readable storage medium. Background Technology

[0002] In the development of intelligent interactive systems, object representation generation faces multiple technical bottlenecks. Current systems, when processing user commands such as "show an image of a rose" or "play rain audio," generally rely on a single data source for representation acquisition. When the target object is missing from the local resource library, the system directly returns a failure result, unable to complete the representation generation task through alternative methods. This rigid acquisition method results in a large number of objects failing to obtain effective representations, especially when dealing with unconventional objects. Simultaneously, the system lacks a continuous learning mechanism for user feedback, repeatedly executing the same process for each identical object, failing to optimize the acquisition strategy based on historical interactions. For example, after a user taps on the generated rose image, the system continues to process subsequent similar requests in the same way, causing a continuous deterioration in user experience. Existing technologies also suffer from limitations in modal support; most systems can only output a single type of representation, unable to simultaneously meet the needs of multimodal interaction such as visual, audio, and tactile feedback. In virtual character interaction scenarios, object representation and action execution often become disconnected, leading to asynchronous character actions and item presentation under the "show item" command. In terms of resource scheduling, the system fails to dynamically adjust its acquisition strategy based on factors such as network conditions and device load, leading to response delays during peak periods. These shortcomings collectively result in a clunky interaction process, limited coverage, and low user satisfaction, severely restricting the practical value of the intelligent system. Existing technologies urgently need improvement to address these issues. Summary of the Invention

[0003] The purpose of this application is to provide an object representation generation method, system, and computer-readable storage medium, which has the advantages of improving the success rate of object representation acquisition, optimizing resource scheduling, enhancing user experience, and continuously improving strategies through feedback learning.

[0004] This application provides a method for generating object representations, the technical solution of which is as follows: An object representation generation method includes the following steps: Retrieve information related to the object; Based on the information, a multi-level priority strategy is used to obtain the representation of the object; Record feedback related to the object and adjust the priority of subsequent representation retrieval based on the feedback.

[0005] Furthermore, this application also proposes a multi-level priority strategy that includes trying multiple representation acquisition methods sequentially.

[0006] Furthermore, this application also proposes multiple representation acquisition methods, including at least one of local matching, network acquisition, or content generation.

[0007] Furthermore, this application also proposes that content generation includes, but is not limited to, at least one of visual content generation, audio content generation, haptic content generation, or motion content generation.

[0008] Furthermore, this application also proposes to apply preset style parameters during content generation to ensure that the generated content is consistent with the style of the object.

[0009] Furthermore, this application also proposes representations including, but not limited to, at least one of visual representation, audio representation, tactile representation, or motion representation.

[0010] Furthermore, this application also proposes to include a representation presented in a manner associated with an object.

[0011] Furthermore, this application also proposes that recording feedback includes recording user feedback.

[0012] Furthermore, this application also proposes that user feedback includes, but is not limited to, at least one of likes, dislikes, or manual corrections.

[0013] Furthermore, this application also proposes to include caching the obtained representations related to the object to avoid repeated retrieval.

[0014] Furthermore, this application also proposes to include actively triggering object representation generation in specific scenarios and soliciting user requests.

[0015] Furthermore, this application also proposes an object representation generation system, comprising: The information acquisition module is used to acquire information related to the object; The representation acquisition module is used to acquire the representation of an object based on information and employing a multi-level priority strategy. The feedback learning module is used to record feedback related to objects and adjust the priority of subsequent representation acquisition based on the feedback.

[0016] Furthermore, this application also proposes that the representation acquisition module be configured to try multiple representation acquisition methods in sequence.

[0017] Furthermore, this application also proposes multiple representation acquisition methods, including at least one of local matching, network acquisition, or content generation.

[0018] Furthermore, this application also proposes that content generation includes, but is not limited to, at least one of a visual content generation unit, an audio content generation unit, a tactile content generation unit, or a motion content generation unit.

[0019] Furthermore, this application also proposes to include a style control module for applying preset style parameters during content generation.

[0020] Furthermore, this application also proposes to include a rendering module for rendering the representation in a manner associated with an object.

[0021] Furthermore, this application also proposes to include a caching module for storing the acquired representation.

[0022] Furthermore, this application also proposes to include an active triggering module for actively triggering object representation generation in specific scenarios.

[0023] Furthermore, this application also proposes a computer-readable storage medium having a computer program stored thereon, which executes the above-described method when executed by a processor.

[0024] As can be seen from the above, the object representation generation method, system, and computer-readable storage medium provided in this application solve the problems of acquisition failure caused by relying on a single data source, lack of learning mechanism, and rigid resource scheduling in the prior art by acquiring object information, adopting a multi-level priority strategy to acquire representation, and adjusting priority based on feedback. It has the advantages of improving the success rate of object representation acquisition, optimizing resource scheduling, enhancing user experience, and continuously improving the strategy through feedback learning.

[0025] Comparative analysis with existing technologies: Regarding object recognition and representation generation: In the prior art, the system can recognize objects, but usually cannot automatically generate representations, or requires users to manually input them into the generation tool; while the solution of this application can realize interaction-driven operation, automatically start a multi-level acquisition process after recognition, and seamlessly connect interaction and representation generation.

[0026] Regarding acquisition strategies: Existing technologies generally rely on a single source (such as an icon library); while the solution proposed in this application can implement a multi-level priority strategy, with a predefined sample library → online search → content generation, progressing step by step to ensure coverage.

[0027] Regarding learning capabilities: existing technologies lack a feedback mechanism; while the solution proposed in this application can achieve closed-loop learning, record user feedback, and dynamically adjust the priority of acquisition strategies.

[0028] Regarding output modality: existing technologies are primarily visual; while the solution of this application can support multimodal output such as visual, audio, and tactile, and can be flexibly expanded.

[0029] Regarding style consistency: Existing technologies may generate content that does not match the context style; while the solution of this application can achieve style control, apply preset style parameters, and ensure that the generated content is consistent with the style of the target object.

[0030] In terms of technical effectiveness: existing technologies are limited to text or single-modal interaction, resulting in low coverage and poor user experience; while the solution proposed in this application can achieve a complete closed loop of "recognition-acquisition-learning-display", supporting multimodal interaction, high coverage, adaptive optimization, and multimodal immersion.

[0031] The above comparison shows that this application is not a simple combination, but rather an organic integration of multi-level strategies, feedback learning, and style control, achieving a general, adaptive, and multimodal object representation generation scheme, which has significant inventiveness. Attached Figure Description

[0032] Several embodiments of this application are described below with reference to the accompanying drawings. It should be noted that the specific structures, modules, steps, parameters, and connections shown in the drawings are preferred embodiments of this application and not limitations on the scope of protection of this application. Those skilled in the art can make various modifications, substitutions, or combinations to the specific details shown in the drawings based on the teachings of this application, and these modified embodiments should still be considered to fall within the scope of protection of this application.

[0033] Figure 1 This application provides a system architecture diagram, which is an exemplary architecture representing the generation system of the object of this application. Each module can be adjusted according to actual applications and does not limit the scope of protection.

[0034] Figure 2 This application provides a flowchart for obtaining a multi-level priority. This diagram illustrates the process of a multi-level priority acquisition strategy. The diagram is merely an example and does not constitute a limitation on the claims.

[0035] Figure 3 This diagram illustrates the process of object representation and action linkage in a smart agent scenario, as provided in this application. Detailed Implementation

[0036] The technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application. Other technologies that may be mentioned in the embodiments can be implemented using existing technology or other patent applications filed by the applicant on the same day, and will not be repeated here. It should be particularly noted that the specific module divisions, process steps, data flow directions, status names, time values, etc., shown in the accompanying drawings are merely illustrative examples and should not constitute a limitation on the scope of protection of the claims of this application. The scope of protection of the claims is determined solely by their wording and should be interpreted in accordance with the overall content of the specification.

[0037] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0038] In the object representation generation process, existing technologies typically rely on a single source to obtain the object representation, such as querying only a local icon library or local audio library for matching. When the target object is missing from local resources, the system cannot switch to other acquisition methods, resulting in representation acquisition failure. Furthermore, the system lacks a feedback recording mechanism; user evaluations of the representation results are not stored and analyzed, making it impossible to optimize the acquisition process for subsequent identical objects. This reduces the success rate of representation acquisition and limits the system's adaptability.

[0039] For example, in a smart assistant application, a user inputs the voice command "play ocean wave sound." The system only matches the local audio resource library. If the corresponding audio file is not found in the library, it directly returns an error message without performing an online search or audio generation operation. The user provides negative feedback to the system's response, but this feedback is not captured and processed by the system. In subsequent interactions, when the user requests the same audio again, the system repeats the same failed process without attempting alternative acquisition methods. As a result, user interaction is interrupted, and the reliability of the system's response decreases.

[0040] If the above issues are not addressed, the system will face an increased failure rate in retrieving representations when handling diverse object representation requests, especially in scenarios with incomplete resource library coverage. The system's inability to dynamically adjust its strategy based on historical interactions leads to a rigid representation retrieval process that cannot adapt to changes in user preferences. In the long run, this will weaken the system's applicability in multimodal interaction environments and reduce user trust in the system.

[0041] In response, this application proposes an object representation generation method, including the following steps: Retrieve information related to the object; Based on the information, a multi-level priority strategy is used to obtain the representation of the object; Record feedback related to the object and adjust the priority of subsequent representation retrieval based on the feedback.

[0042] For ease of understanding, the following explains some key terms in this embodiment: Object representation generation methods refer to a technological process that uses computation to create or acquire a visual, audible, or perceptual representation of a specific object. This method aims to enable systems to provide users with intuitive or interactive content related to objects, based on their needs.

[0043] An object is an entity that is of interest or needs to be represented in a specific context. This entity can be a concrete object (e.g., a flower, a book), an abstract concept (e.g., an emotion, an event), or a state (e.g., the temperature of a machine, the brightness of the environment).

[0044] Information refers to various data or descriptions related to an object. This information can include the object's name, attributes, context, user intent, or any data that helps identify and describe the object. Acquiring information is fundamental to initiating the object representation generation process.

[0045] A multi-level priority strategy refers to a decision-making mechanism that attempts to acquire an object representation according to multiple preset levels or sequences. This strategy aims to improve the success rate and efficiency of representation acquisition by switching between different sources or methods to address the limitations that may exist with a single source.

[0046] Representation refers to the specific form in which an object is presented in a particular modality. This can include, but is not limited to, images, sounds, tactile signals, text descriptions, or sequences of actions. The purpose of representation is to transform abstract object concepts into forms that are perceptible or interactive to the user.

[0047] Feedback refers to the evaluation or response received by a system from external sources (e.g., users, other systems) regarding the quality or suitability of an object representation after the system has acquired and may present such a representation. Feedback is an important basis for the system to learn and optimize itself.

[0048] Priority refers to the order or weight in which different representation acquisition methods or sources are tried in a multi-level priority strategy. By adjusting priorities, the system can optimize its representation acquisition behavior based on historical experience or specific conditions in order to obtain better results.

[0049] This embodiment provides an object representation generation method, which is characterized by the following aspects: First, information related to the object is acquired. This step is the starting point of the entire representation generation process, aiming to provide the system with the necessary context and description of the target object. Specifically, information can be provided by the user through text input, voice commands, or gestures; for example, a user can input "apple" to refer to a fruit. Additionally, information can be obtained through sensor data, environmental monitoring systems, or data interaction with other intelligent systems; for example, the system can obtain information about the "living room temperature" from smart home devices. By acquiring this information, the system can clearly identify the object to be represented and its related attributes, thus providing a clear target for subsequent representation acquisition.

[0050] Secondly, based on the acquired information, a multi-level priority strategy is employed to obtain the object's representation. This step aims to overcome the limitations of a single acquisition method and improve the success rate and efficiency of representation acquisition. In one implementation, the system can pre-define an acquisition order; for example, it might first attempt to find a representation matching the object from a predefined, structured data source. If this attempt fails, the system can automatically switch to a second-level strategy, such as searching through a broader, unstructured data source. If the second-level strategy still fails to yield satisfactory results, the system can further attempt a third-level strategy, such as generating a representation based on the object's description using a computational model. For example, when needing to obtain a representation of "sunset," the system might first attempt to search in an internal image library; if not found, it might try searching in an external image library; if still not found, it might attempt to generate it using an image synthesis algorithm. This multi-level strategy ensures that alternative solutions are available in different situations, thereby improving the robustness of representation acquisition.

[0051] Furthermore, the system records object-related feedback and adjusts the priority of subsequent representation retrieval based on this feedback. This step aims to enable the system to learn and optimize, thereby continuously improving the quality of representation generation and user satisfaction. In one implementation, the system can record the result of each representation retrieval, along with external evaluations related to that result. For example, after generating a representation for an object, the system can receive signals regarding whether the representation meets expectations. These signals can be implicit, such as the user's dwell time on the representation, subsequent actions, etc., or explicit, such as the user directly expressing satisfaction or dissatisfaction through an interactive interface. The system associates and stores this feedback data with the corresponding object, retrieval method, and representation result. Based on this recorded feedback, the system can dynamically adjust a multi-level priority strategy. For example, if a certain retrieval method has frequently received negative feedback in the past, its priority can be reduced when retrieving representations of the same or similar objects in the future, while the priority of well-performing retrieval methods can be increased. In this way, the system can self-optimize based on actual usage, making the representation generation process more intelligent and personalized.

[0052] The following example will provide a more detailed explanation of the above technical solution: Imagine a smart interactive system where user A gives the system the instruction: "Please show me an image of an 'apple'." First, the system performs the "acquiring information related to the object" step. Using speech recognition or text parsing technology, the system identifies the core object as "apple" from user A's instructions and determines that user A's need is to obtain its "image" representation. This information is extracted and used as the target for subsequent representation generation.

[0053] Next, the system executes the step of "obtaining the representation of the object using a multi-level priority strategy based on the information." According to the preset priority strategy, the system begins attempting to obtain an image representation of "apple." For example, the system might first try searching for an image of "apple" in a small local image library. If no suitable image is found in the local library, the system switches to a second-level strategy, such as searching using an external general-purpose image search engine to obtain an image of "apple." If the external search also fails to find satisfactory results, the system might further try a third-level strategy, such as calling a basic image generation model to generate a general image based on the description of "apple." Through this progressive approach, the system aims to ensure that, under all circumstances, it can provide user A with an image representation of "apple" as comprehensively as possible.

[0054] Subsequently, the system executes the step of "recording object-related feedback and adjusting the priority of subsequent representation acquisition based on the feedback." Suppose the system acquires an image of an "apple" through a second-level strategy (an external image search engine) and presents it to user A. After seeing the image, user A might express negative feedback through a simple "dissatisfied" button, or implicitly express dissatisfaction by not taking any further action for an extended period. The system records and associates user A's feedback with the "apple" object and the act of acquiring the representation through the "external image search engine." In the future, when the system needs to acquire a representation for an "apple" or similar object again, it will refer to these recorded feedbacks. For example, because the "external image search engine" received negative feedback when acquiring the "apple" image, the system might lower the priority of the "external image search engine" or increase the priority of other acquisition methods (e.g., the basic image generation model) when acquiring a representation for "apple" in the future, thereby attempting to provide a representation that better meets user A's expectations. In this way, the system can learn from user A's interactions and continuously optimize its representation acquisition strategy to provide more personalized and high-quality services.

[0055] Based on the above examples, the technical concept proposed in this embodiment demonstrates a significant technical contribution. In traditional solutions, when user A requests an image of an "apple," the system may rely solely on a single local image library. If an image of an "apple" is not found in that library, the system cannot provide a valid representation, leading to interrupted interaction or a poor user experience. In contrast, this embodiment introduces a "multi-level priority strategy for acquiring object representations," enabling the system to automatically switch to a second or even third-level strategy when the first-level acquisition fails, such as acquiring from an external search engine or creating it through a generative model. This mechanism greatly improves the success rate and coverage of object representation acquisition, effectively solving the problem of a single acquisition method in traditional solutions.

[0056] Furthermore, traditional solutions generally lack the ability to learn from user feedback. Every time user A requests an image of an "apple," the system repeats the same acquisition process, unable to optimize even if the previously provided representation is unsatisfactory. This embodiment, however, endows the system with adaptive and learning capabilities through the innovative step of "recording object-related feedback and adjusting the priority of subsequent representation acquisition based on the feedback." When user A is dissatisfied with the "apple" image obtained through an external search engine, the system can record this negative feedback and dynamically adjust priorities in subsequent interactions, such as prioritizing other acquisition methods. This closed-loop learning mechanism allows the system to continuously adapt to user preferences and needs over time, thereby continuously improving the quality of representation generation and user satisfaction. Therefore, this embodiment is not simply a combination of existing technologies, but rather an organic integration of multi-level strategies and feedback learning to achieve a universal and adaptive object representation generation scheme, significantly improving the intelligence level and user experience of intelligent interaction systems.

[0057] In some of the solutions described above in this application, a multi-level priority strategy is proposed to flexibly obtain object representations. However, without a clear order of attempts, this process may lead to chaotic strategy execution, wasted resources, or response delays. For example, trying all methods in parallel may cause computational conflicts, or random attempts may reduce acquisition efficiency. Therefore, this application further proposes a multi-level priority strategy that includes sequentially trying multiple representation acquisition methods.

[0058] The multi-level priority strategy refers to a mechanism that sorts and selects different processing or acquisition methods according to preset rules or conditions. Its purpose is to determine an optimal or suboptimal execution path among multiple options based on factors such as efficiency, cost, and success rate. In object representation generation, the multi-level priority strategy ensures that the system can systematically attempt multiple representation acquisition methods, thereby improving acquisition efficiency and success rate. One implementation is based on a preset fixed order, for example, always prioritizing local resources, followed by online search, and finally content generation. This method is simple, direct, and easy to implement and manage. Another implementation is based on dynamically evaluated priorities. For example, the system can adjust the priority order of different acquisition methods in real time based on factors such as historical success rate, resource consumption, user preferences, or current network conditions. For instance, if the success rate of local matching is high, its priority can be dynamically increased; if online acquisition frequently times out, its priority can be dynamically decreased.

[0059] The sequential attempts refer to executing different representation acquisition methods one by one in a pre-set order until a representation is successfully acquired or all methods have been tried. Its purpose is to avoid resource waste and improve efficiency, ensuring that the system does not simultaneously initiate all time-consuming or resource-intensive acquisition processes, but instead prioritizes lower-cost and more efficient methods. One implementation method uses conditional judgment and sequential execution programming logic. For example, it first attempts local matching; if the match is successful, it ends; otherwise, it proceeds to the next step of attempting network acquisition. If successful, it ends; otherwise, it proceeds to the next step of attempting content generation. Another implementation method is based on a state machine model, where each state represents an acquisition method, and the transition between states is determined by the success or failure of the previous method. For example, the initial state is "attempting local matching," if it fails, it transitions to the "attempting network acquisition" state; if that also fails, it transitions to the "attempting content generation" state.

[0060] The various representation acquisition methods refer to the set of different approaches or methods that the system can use to acquire object representations. These methods typically differ in terms of acquisition efficiency, resource consumption, flexibility, and coverage, collectively forming the foundation of the system's ability to acquire object representations. A common implementation may include matching from a predefined local resource library, remote searching via a network connection, or dynamically creating content using a generative model. Furthermore, it may include acquiring data from third-party services by calling specific API interfaces, retrieving from user-defined resource libraries, or capturing data in real time via sensors. These different methods provide the system with a flexible selection space to address various complex object representation needs.

[0061] This application's solution ensures the orderliness and efficiency of the representation acquisition process by employing a multi-level priority strategy and attempting different methods sequentially after acquiring information related to the object. This mechanism allows the system to first utilize efficient, low-resource-consumption acquisition methods, such as matching from a local resource library. If this method fails to obtain the required representation, the system automatically and systematically switches to the next acquisition method, such as online acquisition, according to a preset priority order. If online acquisition still fails, more flexible but potentially resource-intensive methods such as content generation are further attempted. This progressive trial process avoids resource conflicts and waste that could result from simultaneously activating all acquisition methods, and also avoids the inefficiency of random attempts. Furthermore, by providing multiple alternative acquisition methods, the system can seamlessly switch to other methods when one fails, significantly improving the success rate of object representation acquisition. Furthermore, this orderly acquisition strategy, combined with recording feedback and adjusting the priority of subsequent acquisition representations, enables the system to learn from the success or failure of each attempt and dynamically optimize the order or weight of the multi-level priority strategy based on feedback. This allows for continuous self-improvement and adaptation to user preferences, further enhancing overall efficiency and user experience.

[0062] The following example illustrates this. In a smart assistant application, when a user commands "Send me a rose," the system first obtains information related to the object "rose." Then, the system initiates a multi-level priority strategy to obtain a representation of "rose." Specifically, the system first attempts to match within the local icon library to see if a pre-stored rose image exists. If a local match is found, that image is used directly as the representation. If no matching rose image is found in the local icon library, the system then tries the second method: searching online for a rose image. If online searching also fails to find a suitable image, the system continues to try the third method: dynamically generating a rose image using an image generation model. Through this sequential trial strategy, the system ensures that among multiple acquisition methods, it prioritizes the most efficient and resource-saving methods, and gradually upgrades to more complex generation methods when necessary, thus efficiently and reliably obtaining the object's representation.

[0063] Through the above technical solution, this application effectively solves the problems of chaotic strategy execution, resource waste, or response delays in traditional methods. By clearly defining the sequential attempt order, the system can efficiently utilize resources, prioritizing lower-cost and more efficient methods, and avoiding unnecessary computational overhead and network requests. Simultaneously, the introduction of multiple representation acquisition methods ensures that there are alternative solutions when one method fails, thereby significantly improving the success rate of object representation acquisition. This orderly and flexible strategy makes the system more stable and efficient in acquiring object representations, and provides a solid foundation for subsequent feedback learning and optimization.

[0064] In some of the solutions described above in this application, a multi-level priority strategy is proposed, including multiple representation acquisition methods to improve the coverage and success rate of object representation. However, in its implementation, if the specific types and scope of these acquisition methods are not defined, it may lead to ambiguity in method selection, low resource scheduling efficiency, and an inability to fully cover resource availability and personalized needs in different scenarios, thereby affecting the feasibility and actual effect of the strategy. Therefore, this application further proposes multiple representation acquisition methods, including at least one of local matching, network acquisition, or content generation.

[0065] Specifically, local matching refers to the system searching for a representation that matches the target object in a pre-stored local resource library. This local resource library can be a structured database storing object identifiers and their corresponding representation file paths; it can also be a file system directory organizing various representation files according to specific naming rules or metadata. When the system receives object information, it first attempts to quickly locate and retrieve the corresponding representation in this local resource library using methods such as keyword matching, hash value comparison, or feature vector retrieval. As another implementation approach, the system can pre-cache representations of frequently accessed or used objects, forming a personalized local cache library. When a request arrives, it prioritizes searching in this cache library.

[0066] Network-based retrieval refers to the system obtaining the required representation of an object by connecting to external resources via a network when local matching fails to find the object's representation. This can be achieved by calling public internet search engine APIs, using object information as query criteria to retrieve relevant resources such as images, audio, or video; or by connecting to specific online content platforms, cloud services, or third-party databases. These platforms or services specialize in providing various types of object representations, such as online image libraries, sound effect libraries, and 3D model libraries, with the system requesting and downloading data through the corresponding interfaces.

[0067] Content generation refers to the system dynamically creating object representations using generative models when local matching and online acquisition cannot meet the requirements, or when highly personalized and customized representations are needed. This can be achieved by calling pre-trained deep learning models, such as Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), or Diffusion Models, to synthesize content in modalities such as visual images, audio clips, tactile signals, or action sequences based on the object's textual description, attribute features, or other input information. Alternatively, it can utilize large language models (LLMs) or multimodal large models, combined with context and object descriptions, to directly output or assist in the generation of various modal representations, such as generating image prompts, audio synthesis parameters, or action sequence instructions.

[0068] This application's solution, by enumerating specific representation acquisition methods within a multi-level priority strategy—namely, local matching, network acquisition, and content generation—enables the entire object representation generation method to possess high robustness and flexibility. After acquiring object-related information, the system attempts these acquisition methods sequentially according to a preset priority order. First, the system attempts local matching, leveraging the fast response and low-cost advantages of local resources, suitable for high-frequency or pre-cached object representation needs. If local matching fails, the system switches to network acquisition, greatly expanding the representation coverage and compensating for the limitations of local resources by accessing a wider range of external network resources. When neither local nor network methods can provide a satisfactory representation, or when unique, personalized content needs to be generated, the system initiates a content generation mechanism, dynamically creating the required representation using an advanced generation model. This progressively advancing trial mechanism ensures that the system can effectively acquire or generate object representations under varying resource availability, network conditions, and personalized requirements. Furthermore, by combining the mechanism of recording object-related feedback and adjusting the priority of subsequent representation acquisition, the system can dynamically optimize the weight or attempt order of each method based on the user's satisfaction with the results generated by different acquisition methods, thereby forming an adaptive and continuously learning closed loop, which further improves the success rate of representation acquisition and user experience.

[0069] As a specific implementation, in scenarios where a smart assistant interacts with a user, when the user issues the command "Show me a rose," the system first identifies the target object as "rose." Then, the system initiates the representation acquisition process. First, the system attempts to perform a local match in its locally stored icon library or image cache to check if a pre-stored "rose" image exists. If the local match is successful, the system will directly use that local image for presentation. If the local match fails to find a suitable "rose" image, the system automatically switches to the second step, namely, online acquisition. At this point, the system may send a query request for "rose" to an internet image search engine or access an online plant image database to obtain relevant image resources. If online acquisition also fails, or if the user is not satisfied with the generic image obtained online, the system will proceed to the third step, content generation. The system will call a pre-trained image generation model (e.g., a text-based image model based on a diffusion model) to dynamically generate a "rose" image using "rose" as the input prompt. During the generation process, the system can also apply preset style parameters to ensure that the generated image maintains consistency with the overall visual style of the smart assistant. By trying each option in turn, the system can maximize the chances of providing the user with a representation of a "rose".

[0070] Through the above technical solutions, this application lists specific types of representation acquisition methods in a multi-level priority strategy, effectively solving the problems of ambiguous method selection, low resource scheduling efficiency, and inability to fully cover resource availability and personalized needs in different scenarios in traditional solutions. Specifically, the introduction of local matching enables the system to prioritize the use of local resources, significantly improving response speed and acquisition efficiency, and reducing dependence on external networks. Network acquisition, as a supplement, greatly expands the range of available representations, ensuring that rich representations can still be obtained when local resources are insufficient. Content generation provides the final guarantee, enabling dynamic creation of representations according to needs even without any pre-stored or network-available resources, especially meeting the needs for personalized and customized content. This combination not only improves the success rate and coverage of representation acquisition, but also optimizes resource utilization through priority scheduling, allowing the system to flexibly select the most suitable acquisition method according to the actual situation. Combined with a feedback learning mechanism, the system can dynamically adjust the priority of these methods based on user satisfaction with the results generated by different acquisition methods, thereby continuously optimizing the representation acquisition strategy and achieving a more intelligent and personalized interactive experience.

[0071] Traditional solutions have proposed content generation as one of the representation acquisition methods to dynamically generate object representations. However, in this process, content generation may be limited to a single modality and cannot cover other modalities such as audio, touch, or motion. This results in the inability to provide comprehensive representations in scenarios that require multi-sensory interaction, thus limiting the richness and adaptability of the user experience.

[0072] In this regard, this application further proposes content generation including but not limited to at least one of visual content generation, audio content generation, haptic content generation, or motion content generation.

[0073] Visual content generation refers to the creation of visual representations such as images, videos, and 3D models through computational methods. This can utilize deep learning models, such as Generative Adversarial Networks (GANs) or diffusion models, to generate realistic or stylized images based on text descriptions, sketches, or other inputs. Additionally, 3D rendering techniques can be used to construct virtual 3D objects or environments based on object attributes and scene parameters. Audio content generation refers to the synthesis of auditory representations such as sound, speech, or music through algorithms. This can be achieved by synthesizing natural language speech using text-to-speech (TTS) technology, or by using sound synthesizers to generate environmental sound effects (such as rain or ocean waves) or specific timbres. Another approach is procedural audio generation, dynamically generating complex soundscapes or sound effects based on physical models or abstract rules. Haptic content generation refers to the generation of tactile feedback such as vibration, pressure, or texture through technological means. This can be achieved through tactile rendering algorithms, converting the physical properties of virtual objects (such as roughness, softness, and elasticity) into specific vibration patterns or force feedback signals, and then transmitting them to the user through tactile devices (such as vibration motors or force feedback gloves). Electromechanical actuators can also be used to simulate the textures or surface properties of different materials. Motion content generation refers to generating a series of motion sequences or animations for virtual characters, robots, or other movable entities. This can be achieved using inverse kinematics (IK) or forward kinematics (FK) algorithms to drive virtual skeletal animation based on the target pose or motion trajectory. Furthermore, deep learning-based motion synthesis techniques, such as the Transformer model, can be used to generate smooth and natural motion sequences from high-level instructions or text descriptions.

[0074] In some of the above embodiments, after acquiring information related to an object, the system employs a multi-level priority strategy to obtain the object's representation, one of which is content generation. This application further clarifies that when the system needs to obtain an object representation through content generation, this content generation is no longer limited to a single modality but can encompass multiple modalities such as visual content generation, audio content generation, haptic content generation, or motion content generation. Specifically, when local matching and online acquisition fail to provide the required object representation, or when a personalized, specific modality representation needs to be generated, the system will initiate a content generation mechanism. At this time, the system will intelligently select or combine visual content generation, audio content generation, haptic content generation, or motion content generation based on the current interaction scenario, user needs, and the characteristics of the object. For example, if a user needs an image of an object, the system will generate visual content; if an ambient sound effect is needed, audio content will be generated; if a certain tactile sensation needs to be simulated, haptic content will be generated; if a virtual character needs to perform a specific action, motion content will be generated. This multimodal content generation capability, combined with a multi-level priority strategy, ensures that even without preset resources, the system can dynamically create representations that meet user needs and scenario characteristics, greatly improving the success rate and diversity of representation acquisition. In this way, the system can provide a more comprehensive and richer sensory experience, effectively overcoming the limitations of traditional content generation methods that rely on a single modality.

[0075] The following example illustrates this. In a virtual reality (VR) shopping scenario, a user might express, "I want a red plush toy." Upon receiving this information, the system first attempts to match it in its local resource library. If no match is found, it searches online for relevant images or models. If neither method provides a representation that perfectly matches the user's description (especially the tactile attribute of "plush"), the system initiates content generation. At this point, the system simultaneously generates visual and tactile content. The visual content generation module uses a diffusion model to generate a 3D model or image of a red plush toy, ensuring its visual features match the description. Simultaneously, the tactile content generation module generates a vibration pattern or force feedback signal simulating a soft, fluffy feel based on the "plush" attribute, and outputs it through the user's VR controller or haptic gloves. Furthermore, if needed, the system can also generate audio content, such as upbeat background music related to the toy's theme. Through this multimodal content generation, the system can provide users with a virtual toy representation that they can see, touch, and even hear, thus achieving a more immersive and realistic interactive experience.

[0076] Through the above technical solution, this application effectively solves the problem of single modality in traditional content generation methods. By extending content generation to multiple modalities such as visual, audio, tactile, and motion, the system can dynamically generate diverse sensory representations based on different interaction needs and scene contexts. This not only greatly improves the coverage and success rate of object representation, ensuring appropriate representations can be provided even in the absence of pre-built resources, but also significantly enhances the richness and immersion of the user experience. For example, in scenarios requiring multi-sensory interaction, the system is no longer limited to a single visual or auditory output, but can provide tactile feedback or motion linkage, making human-computer interaction more natural and vivid. This combination of multi-modal content generation capability and multi-level priority strategy enables the system to respond to user needs more flexibly and intelligently, thereby providing more comprehensive and personalized object representations.

[0077] In some of the embodiments described above in this application, content generation is proposed to obtain object representations. However, during its implementation, the generated content may be inconsistent with the style of the target object, resulting in content that is inconsistent with the context and lacks personalization, thus affecting user experience and interactive immersion. To address this, this application further proposes applying preset style parameters during content generation to ensure that the generated content is consistent with the object's style.

[0078] "Content generation time" refers to the process during which the system dynamically creates object representations, i.e., during the execution of the generative model or in a closely related real-time phase. This timing ensures that style control directly influences the output of the generation process, rather than acting as a post-processing fix, thus avoiding additional processing delays and improving the effectiveness of control. "Preset style parameters" refer to predefined or configured settings used to guide the content generation model in outputting specific style characteristics. These parameters can vary depending on the modality and application scenario. For example, for visual content generation, style parameters may include specific color schemes, texture types, art styles (such as cartoon, realistic, watercolor), or lighting conditions; for audio content generation, style parameters may involve timbre, rhythm, emotional tone, or instrument type; for haptic content generation, style parameters may define vibration patterns, intensity, frequency, or simulated material textures (such as rough, soft); for motion content generation, style parameters may include the smoothness, speed, expressiveness of the motion, or behavioral characteristics of a specific character. These parameters can be provided to the generative model in the form of numerical values, text descriptions, reference samples, or learned embedding vectors. "Ensuring that the generated content is consistent with the object's style" means applying the aforementioned preset style parameters to ensure that the generated object representation maintains harmony and consistency with the inherent style of the target object, user preferences, or the overall style of the current interactive scenario in terms of visual, auditory, tactile, or behavioral aspects. This consistency aims to eliminate the abruptness caused by style mismatch and enhance user acceptance and immersion in the generated content.

[0079] This application's solution addresses the issue of inconsistent styles between generated content and objects by deeply integrating style control mechanisms into the content generation process. Specifically, when the system determines, based on acquired object-related information and a multi-level priority strategy, that content generation is necessary to obtain the object's representation, preset style parameters are applied simultaneously as the content generation module initiates the generation process. These style parameters serve as guides or constraints for the generation model, directly influencing its output. For example, when creating a visual representation, the generation model adjusts its internal texture, color, and composition logic based on the input style parameters; when generating an audio representation, it adjusts the synthesis algorithm based on parameters such as timbre and rhythm; and when generating tactile or motion representations, it adjusts vibration patterns or motion trajectories according to the corresponding style parameters. This approach of style control during the generation stage allows the final output object representation to naturally integrate into the context of the target object or the personalized experience desired by the user, thus avoiding the complexity and limitations of later adjustments. In this way, even when local matching and online acquisition cannot meet the requirements and dynamic generation of representation is necessary, the system can ensure that the generated representation is not only usable, but also highly consistent with the style of the object, greatly improving the consistency and immersion of the user experience.

[0080] The following example illustrates this. Suppose in a smart assistant application, a user commands the assistant, "Show me a watercolor-style rose." The system first obtains information related to the object "rose" and the user-specified style parameter "watercolor style." Based on a multi-level priority strategy, the system might first attempt to search its local resource library for a rose image matching the "watercolor style." If not, it will attempt to retrieve one online. If online retrieval also fails to find an image perfectly matching the style requirements, the system will activate the content generation module to dynamically generate a visual representation of the rose. During content generation, the system inputs "watercolor style" as a preset style parameter into the image generation model. This model then generates a rose image with characteristics of a watercolor painting based on these style parameters. For example, the edges of the image will exhibit the characteristic watercolor wash effect, the color transitions will be soft, and the overall visual effect will conform to the artistic style of a watercolor painting. In this way, the system ensures that the generated rose image is not only a rose, but also that its visual style is consistent with the user-specified "watercolor style."

[0081] Through the above technical solution, this application effectively solves the problem of style inconsistency that may occur during content generation. By applying preset style parameters during content generation, the system can ensure that the generated visual, audio, tactile, or motion representations are highly coordinated with the inherent style of the target object, user preferences, or the overall style of the current scene. This significantly improves the coordination and personalization of the generated content, avoids the abruptness caused by style mismatch, and thus greatly enhances the immersion and satisfaction of the user experience. Under a multi-level priority strategy, when the system needs to obtain object representations through content generation, this solution ensures that even dynamically generated representations can meet the requirements of high quality and high consistency, further improving the practicality and user acceptance of the entire object representation generation method.

[0082] In some of the solutions mentioned above in this application, representations are proposed to obtain representations of objects. However, in this process, the representation may be limited to a single modality and cannot support multimodal representations such as visual, audio, tactile, or motion, resulting in inflexible interaction and inability to adapt to complex scenarios.

[0083] In this regard, this application further proposes that the representation includes, but is not limited to, at least one of visual representation, audio representation, tactile representation or motion representation.

[0084] In this application, "representation" refers to any form of data or signal that describes or presents a specific object. The concept encompasses various perceptible or manipulable manifestations of an object in the digital or physical world. For example, a representation can be a data structure for storing object attributes; a file format for carrying the object's specific content; or a physical signal for driving external devices to generate perception. "Visual representation" refers to a form of conveying object information through the visual channel. Its implementation may include, but is not limited to: generating image files for display on a screen; generating video streams for demonstrating the object's dynamic changes; generating 3D model data for constructing a three-dimensional image of the object in a virtual environment; or generating a text description that clearly outlines the object's visual characteristics. "Audio representation" refers to a form of conveying object information through the auditory channel. Its implementation may include, but is not limited to: generating audio files for playback through speakers; synthesizing speech or music with specific timbres to meet personalized needs; or generating a text description that accurately conveys the object's auditory characteristics. "Tactile representation" refers to a form of conveying object information through the tactile channel, designed to simulate the sensation of physical contact. The implementation methods may include, but are not limited to: generating specific vibration patterns to simulate rough, soft, or other tactile sensations through haptic feedback devices; generating force feedback signals to simulate the weight or resistance of an object through force feedback devices; or generating pressure distribution data to drive multi-touch displays. "Motion representation" refers to a form describing the motion or behavioral sequence of an object. Its implementation methods may include, but are not limited to: generating animation sequence data to drive virtual characters or robots to perform specific actions; generating motion capture data streams to reproduce real-world actions in real time; or generating a set of robot control commands to precisely control the motion trajectory of robotic arms or other automated equipment. Furthermore, the phrase "including but not limited to at least one of..." emphasizes the openness and flexibility of the representation modalities supported by this application. It means that the system is not limited to the visual, audio, haptic, or motion representations listed above, and can be extended to support other modal representations according to actual needs and technological developments. At the same time, "at least one" ensures that the system can provide at least one effective object representation under any circumstances, even if multimodal combinations cannot be provided due to resource or contextual constraints.

[0085] This application addresses the problems of single representation modality and inflexible interaction in traditional methods by expanding the representation type to include, but is not limited to, at least one of visual, audio, tactile, or action representations. Specifically, after acquiring object-related information, the system is no longer limited to generating a single-modal representation but can flexibly select or combine multiple modal representations based on object characteristics, user needs, or the context of the current scene. For example, for a "rose" object, the system can generate its visual representation; for a "wave sound" object, it can generate its audio representation; for a "cat's fur" object, it can generate its tactile representation; and for a "victory sign" object, it can generate its action representation. This multimodal representation capability is closely integrated with the multi-level priority strategy and feedback learning mechanism in the underlying method. Under the multi-level priority strategy, the system can first attempt to acquire a representation of a certain modality (e.g., visual representation). If this fails or does not meet the requirements, it can switch to acquiring representations of other modalities (e.g., audio or tactile representations), or attempt to acquire a combination of multiple modal representations. For example, when a user requests a "red plush toy," the system can prioritize generating a visual representation while also attempting to generate a tactile representation to provide a richer interactive experience. Furthermore, a feedback learning mechanism can further optimize the acquisition of multimodal representations. The system records user feedback on different modal representations and adjusts the priority of subsequent representation acquisition based on this feedback. For instance, if a user frequently taps on the visual representation of a certain object, the system may prioritize generating its audio or tactile representation when encountering that object again, or adjust the generation parameters of the visual representation. This mechanism allows the system to continuously learn user preferences, optimizing not only the representation acquisition method but also the modal selection of representations, thereby providing a more personalized and intelligent interactive experience. In this way, the method of this application can adapt to a wider range of interaction scenarios, providing a more natural and immersive user experience, effectively overcoming the limitations of single-modal representations.

[0086] The following is a concrete example. In a smart assistant application, when a user interacts with the smart assistant, the system needs to provide the user with representations of objects. As a specific implementation, when the user says, "Show me a rose," the system first obtains information related to the object "rose." Next, the system attempts to obtain a representation of "rose" based on a multi-level priority strategy. Since the user explicitly requests to "see," the system prioritizes obtaining a visual representation. This might include searching for an image of a rose in a local library or searching online for a picture of a rose. If these methods fail to obtain a satisfactory visual representation, the system will call a content generation model to generate an image of a rose. Furthermore, if the user says, "Play the sound of waves," after obtaining information about the object "wave sound," the system prioritizes obtaining an audio representation. This might include matching wave sound audio in a local audio library or obtaining wave sound audio online. If these methods fail to obtain a satisfactory audio representation, the system will call an audio generation model to generate the wave sound audio. For another example, in a virtual reality (VR) experience, when a user tries to "pet a virtual cat," the system acquires information about the object's "cat fur." ​​At this point, the system will prioritize acquiring tactile representation. This might be done by calling a tactile signal generation model to generate corresponding vibration patterns or force feedback signals based on the characteristics of the "cat fur," and then outputting them to the user through a VR controller or haptic gloves to simulate a realistic tactile sensation. Furthermore, in a virtual character interaction scenario, when a user says "make a victory sign" to the agent, the system acquires information about the object's "victory sign." The system will prioritize acquiring motion representation. This might include searching for animation data of the "victory V-sign" in a predefined motion library or obtaining relevant motion animations through an online search. If these methods fail to obtain a satisfactory motion representation, the system will call a motion generation model to generate a natural victory sign animation. As can be seen from the above examples, the method of this application can flexibly generate representations of different modalities based on the nature of the object and the user's needs, thereby providing a richer and more diverse interactive experience.

[0087] Through the above technical solution, this application effectively solves the problems of traditional methods, such as single representation modality, inflexible interaction, and inability to adapt to complex scenarios. Specifically, by supporting visual representation, audio representation, tactile representation, or action representation, the system can provide the most suitable or richest representation form based on the characteristics of the object, user intent, and current context. This allows interaction to extend beyond single visual or auditory feedback to more immersive modalities such as touch and action, greatly improving the flexibility and naturalness of the interaction. Furthermore, the advantages of this multimodal support are even more significant when combined with the multi-level priority strategy and feedback learning mechanism in the basic method. The multi-level priority strategy can be dynamically adjusted according to the difficulty of acquiring different modalities, resource consumption, or user preferences, ensuring intelligent selection and switching between multiple modalities. The feedback learning mechanism can learn from user feedback on different modal representations, continuously optimizing the system's strategy for selecting which modal representation to use in specific scenarios, thereby achieving more personalized and intelligent object representation generation. Ultimately, the method of this application can provide users with a high-coverage, high-success-rate, continuously optimized, and multimodal interactive experience, significantly enhancing the intelligence of the system and user satisfaction.

[0088] In some of the solutions mentioned above in this application, a representation of an object is proposed to provide interactive content. However, in this process, the representation presentation method may not be associated with the object, resulting in awkward interaction and a disconnect from the user experience.

[0089] In this regard, this application further proposes that, after obtaining the representation of the object, the representation shall also be presented in a manner associated with the object.

[0090] "Representation" refers to outputting the acquired object representation in a way that is perceptible to the user. Its function is to transform abstract or generated information into a concrete sensory experience, thereby achieving human-computer interaction. Specifically, for visual representation, it can be displayed on a screen, projection device, or augmented / virtual reality headset; for audio representation, it can be played through speakers, headphones, or bone conduction devices; for tactile representation, it can generate corresponding physical stimuli through tactile feedback devices (such as vibration motors or force feedback gloves); and for motion representation, it can drive virtual characters to perform animated demonstrations or control physical robots to perform specific actions.

[0091] "In an object-related manner" means that when presenting the representation, a presentation format closely integrated with the object itself or its environment is chosen based on the object's attributes, type, context, or user intent. This relevance aims to avoid the isolated presentation of the representation, making it a natural extension of the object within the interactive environment. For example, when the object is a virtual item, its visual representation can be designed to be "held" or "placed" in a virtual scene by a virtual character; when the object is an abstract concept, its audio or visual representation can be integrated as a background element into the relevant interactive context; when the object is a simulation of a physical entity (such as "cat fur"), its tactile representation can be generated synchronously when the user "touches" the virtual object through a haptic device worn on the user's hand, simulating a real tactile sensation; when the object is a specific action, its representation can be an animation of a virtual character performing the action, or the physical process of a robot performing the action.

[0092] The solution of this application first acquires information related to the object, and based on this information, uses a multi-level priority strategy to acquire the representation of the object. After successfully acquiring the representation, this application further presents the representation in a manner associated with the object. This means that the system does not simply output the acquired visual, audio, tactile, or motion representations independently, but intelligently analyzes the characteristics of the object, the context of the user's request, and the modality of the acquired representation, thereby selecting the most appropriate presentation method that can form an organic whole with the object. For example, if the acquired representation is an image of a virtual rose, and a virtual assistant exists in the interaction scene, the system can drive the virtual assistant to perform a "picking up" or "delivering" action, while displaying the rose image in the assistant's hand, making the user feel that the virtual assistant has actually delivered a flower. If the acquired representation is an audio representation of the sound of waves, the system will play it through a speaker to simulate environmental sound effects. If the acquired representation is a tactile representation simulating a soft touch, the system will provide corresponding tactile feedback synchronously when the user "contacts" the virtual object through a tactile device worn by the user. This closely integrated presentation method ensures that the representation perceived by the user is no longer independent information detached from the object, but rather a natural extension of the object within the interactive environment, thus significantly enhancing the realism, immersion, and coherence of the interaction. Furthermore, the effectiveness of this integrated presentation method can be incorporated into a feedback mechanism that records information related to the object. User satisfaction or dissatisfaction with the presentation method can be recorded and used to adjust subsequent strategies for acquiring and presenting representations, forming a closed loop for continuous optimization of the user experience.

[0093] The following is a concrete example. In a smart home system, a user says to the smart speaker, "Please play a sound of birdsong in a forest." The system first obtains information related to the object "birdsong" and identifies that the user's intention is to play audio.

[0094] Next, the system employs a multi-level priority strategy to obtain the audio representation of the "birdsong". It may first search in the local audio library; if found, it uses it directly; if not found, it attempts to obtain it via the internet; if obtaining via the internet also fails, or a more personalized sound is needed, it calls the audio generation model to generate a birdsong. Assume the system successfully obtains a high-quality birdsong audio file.

[0095] The system then presents the representation in an object-related manner. In this case, instead of simply playing a standalone audio file, the system plays the birdsong through the smart speaker's speaker and adjusts the volume and reverberation based on the room environment, making the user feel as if the birdsong is truly coming from a corner of the room, creating an atmosphere of being in a forest.

[0096] During this process, the system also records feedback related to the object. If the user is satisfied with the playback effect, the system records positive feedback and prioritizes this environmentally friendly playback method in future similar requests. If the user feedback is that the sound is too abrupt or unnatural, the system records negative feedback and adjusts the audio playback parameters or tries other presentation methods in the future.

[0097] Through the above technical solution, this application effectively solves the problem of insufficient relevance between the representation and the object, leading to stiff interaction and a fragmented user experience. After obtaining the representation of the object, this application no longer simply displays the representation independently, but selects a presentation method that is highly integrated with the object based on the object's type, attributes, and context. This associative presentation method ensures that the representation perceived by the user is no longer isolated information, but a natural extension of the object itself within the interactive environment, greatly enhancing the realism, immersion, and coherence of the interaction. Users can understand and experience the object more naturally, thereby significantly improving the overall user experience and avoiding the sense of fragmentation caused by the disconnect between the representation and the object.

[0098] In some of the solutions mentioned above in this application, it is proposed to record feedback to adjust the priority of subsequent representation acquisition based on the feedback. However, if user feedback is not recorded in this process, the feedback source may not be directly relevant, making it impossible to effectively capture user preferences and true intentions, thereby affecting the accuracy of priority adjustment and the effect of personalized optimization.

[0099] In this regard, this application further proposes recording feedback, including recording user feedback.

[0100] "Recording user feedback" refers to the system collecting users' direct evaluations or actions regarding generated or presented object representations. One implementation method is to provide explicit user interface elements, such as displaying "satisfied" or "unsatisfied" option buttons after the object representation is presented, or providing a rating slider where users can express their evaluation by clicking or dragging. The system then records these explicit user inputs as feedback. Another implementation method is to analyze implicit user behavior patterns to infer user feedback. For example, when a user views an object representation for an extended period, replays it, or shares it with other users, the system can interpret this as positive feedback; conversely, if a user quickly closes, skips, or deletes an object representation, it may be interpreted as negative feedback, and the system records this behavioral data.

[0101] This application's solution ensures the directness and effectiveness of feedback data by prioritizing user feedback. Specifically, after acquiring information related to the object, the system uses a multi-level priority strategy to obtain the object's representation. After the representation is presented to the user, the system actively or passively collects the user's direct evaluation or operational behavior regarding the representation. This user feedback is accurately recorded and used as the basis for adjusting the priority of subsequent representation acquisition. In this way, the system can directly perceive user satisfaction or preferences, enabling the feedback learning module to optimize its multi-level priority strategy based on the user's true intentions. For example, the system lowers the priority of acquisition methods that users are generally dissatisfied with, while raising the priority of acquisition methods that users are generally satisfied with. This mechanism avoids reliance on indirect or vague feedback information, ensuring the accuracy and personalization of priority adjustments, thus making the entire object representation generation process more intelligent and user-friendly.

[0102] The following is a concrete example. Suppose an intelligent assistant system, after a user commands "Show me a rose," generates an image of a rose using content generation and presents it to the user. Simultaneously, the system provides a simple "Like" button and a "Dislike" button on the user interface. If the user clicks the "Like" button, the system records this explicit positive user feedback and associates it with the "rose" object and the "content generation" acquisition method. Conversely, if the user clicks the "Dislike" button, the system records negative user feedback. In another scenario, if the system generates an audio representation of "waves," and the user manually stops playback after a few seconds and switches to another task, the system interprets this implicit behavior of quickly interrupting playback as negative user feedback and records it. This recorded user feedback data will directly guide the system in adjusting its multi-level priority strategy to better meet user preferences when encountering similar requests in the future.

[0103] Through the above technical solution, the system can directly obtain users' real evaluations and preferences, thereby ensuring that the feedback data used to adjust the priority of subsequent representation acquisition is highly relevant and accurate. This significantly improves the efficiency and accuracy of the feedback learning module, enabling the system to more effectively adapt to users' personalized needs, continuously optimize the generation quality of object representations and user experience, and avoid optimization biases caused by unclear feedback sources.

[0104] In some of the solutions described above in this application, user feedback is recorded to adjust the priority of subsequent object representation based on the feedback. However, in this process, the types of user feedback may not be specific or diverse enough, making it impossible to accurately capture user preferences, errors, or correction needs, thereby affecting the accuracy and optimization effect of priority adjustment.

[0105] In this regard, this application further proposes user feedback including but not limited to at least one of likes, dislikes, or manual corrections.

[0106] User feedback refers to the evaluation, correction, or interactive behavior of users towards the object representation generated or provided by the system. It can be implemented through specific interactive elements in the graphical user interface (GUI) (such as buttons and sliders), or through voice commands or gesture recognition. "Like" indicates that the user expresses satisfaction or approval with the object representation. This can be achieved by clicking the "Like" or "Thumbs Up" button on the interface, or by saying words like "good" or "like" via voice command. "Dislike" indicates that the user expresses dissatisfaction or disapproval of the object representation. This can be achieved by clicking the "Dislike" or "Dislike" button on the interface, or by saying words like "bad" or "dislike" via voice command. Manual correction allows users to directly modify or provide a more accurate object representation. For example, users can modify the description of the object representation through a text input box, or upload a new representation file; or users can directly modify the generated visual content through drag-and-drop, editing tools, or record new audio or actions.

[0107] This application's solution significantly enhances the functionality of the feedback learning module by introducing specific user feedback types. After the system acquires the representation of an object based on relevant information and employs a multi-level priority strategy, it presents this representation to the user. At this point, the user can provide explicit feedback, such as liking, disliking, or manual correction. When a user likes, the system recognizes that the current object representation meets the user's expectations, thus prioritizing the acquisition method used to generate the representation (such as local matching, online acquisition, or content generation) to encourage future attempts at this successful strategy when encountering similar objects. Conversely, when a user dislikes, the system recognizes that the current object representation does not meet the user's expectations, thus lowering the priority of the acquisition method used to generate the representation, or even marking it as unrecommended, prompting the system to try other acquisition methods in the future. More importantly, when a user manually corrects, the system not only identifies problems with the representation but also directly obtains the correct or better representation information expected by the user. This specific correction information can be used to update the system's knowledge base or directly as new training data to optimize the content generation model, thereby more accurately adjusting the priority of subsequent representation acquisition. In this way, the method of recording feedback related to the object and adjusting the priority of subsequent acquisition of the representation based on the feedback can obtain more instructive fine-grained information from vague "good" or "bad" feedback, making priority adjustment more accurate and effective, thereby continuously optimizing the object representation acquisition process.

[0108] The following example illustrates this. Suppose a user requests "Give me a picture of a retro sports car" through a smart assistant. The system first obtains the object information "retro sports car". The representation acquisition module will attempt to acquire representations sequentially according to a multi-level priority strategy. For example, the system might first try local matching but fail to find a suitable image; then it might try online acquisition and successfully obtain an image of a regular sports car, which is then presented to the user. If the user finds that this image is not the "retro" style they expected, they can choose to click the "dislike" button on the interface. The system records this negative feedback and lowers the priority of online acquisition for objects like "retro sports cars". Subsequently, the user might further correct the image manually through a text input box, entering "I want a classic sports car from the 1950s, such as a Cadillac Eldorado". After receiving this manual correction, the system will call the content generation model to generate a retro sports car image that meets the requirements based on the new description. When this image is presented to the user, if the user is satisfied, they can click the "like" button. The system records this positive feedback and increases the priority of content generation for objects like "retro sports cars". In this way, the system can learn from the user's specific feedback and continuously optimize its object representation acquisition strategy.

[0109] Through the above technical solution, this application addresses the problem of insufficient specificity or diversity in user feedback types in traditional feedback mechanisms. By introducing three preferred user feedback types—likes, dislikes, and manual corrections—the system can obtain more granular information with greater guidance. Likes directly reflect users' positive approval of the generated representation, enabling the system to identify successful cases and prioritize the reuse of corresponding acquisition methods; dislikes clearly indicate user dissatisfaction or errors, allowing the system to promptly reduce the priority of relevant acquisition methods or try alternative strategies; manual corrections allow users to proactively provide correction information, ensuring the system can directly obtain accurate data and quickly adjust its strategies. These specific types of feedback collectively enrich the dimensions of feedback data, resolve ambiguities that may arise from generalized feedback, and enable the feedback learning module to more accurately capture user preferences, errors, or correction needs, thereby more effectively supporting priority adjustments. This not only improves the accuracy and optimization effect of priority adjustments but also allows the object representation generation method to adapt to user needs more quickly and intelligently, ultimately significantly improving the overall intelligence of the interaction and the user experience.

[0110] In some of the solutions described above in this application, a multi-level priority strategy is proposed to obtain object representations. However, in this process, the system needs to re-execute the acquisition process every time it encounters the same object, resulting in wasted resources, response latency, and low efficiency. To address this, this application further proposes to associate the acquired representations with object-related information and cache them to avoid duplicate acquisition.

[0111] This scheme caches the associated information of the acquired representations with the object, aiming to achieve storage and fast retrieval of acquired representations by establishing a mapping relationship between representations and object information. Specifically, a data structure can be constructed, such as a hash table or key-value pair storage, where the key is the object's unique identifier or related information (such as object name, attribute description, context, etc.), and the value is the corresponding representation data. When the system obtains the representation of an object, it stores the representation and its associated information together in this data structure. Alternatively, a database management system (such as a relational or non-relational database) can be used to store this associated data. In the database, a table can be designed containing object ID, object-related information fields, and representation data fields, thereby achieving persistent associated storage of representations and object information. Avoiding duplicate retrieval is the direct effect and purpose of "associative caching," which reduces unnecessary computation and network requests by utilizing cached data. Each time an object representation is needed, the system first checks if the representation exists in the cache. If it exists, it is read directly from the cache and used, thus avoiding the need to execute the multi-level priority strategy retrieval process again. In addition, by setting cache invalidation mechanisms (such as time-based, size-based, or access frequency-based eviction policies), the validity and timely updates of cached data can be ensured, while avoiding problems caused by expired or inaccurate cached data. This ensures data freshness while minimizing duplicate retrieval.

[0112] The proposed solution associates the acquired representation with object-related information in a cache. This allows the system to query the cache first after acquiring object-related information, before executing a multi-level priority strategy to obtain the object's representation. If the object's representation already exists in the cache, the system retrieves it directly, skipping the multi-level priority strategy and significantly reducing response time and saving computational resources. The system only initiates the multi-level priority strategy to obtain the representation when it is not present in the cache or when the cached data is invalid. Once the representation is successfully obtained through the multi-level priority strategy, it is immediately written to the cache along with the object-related information. Furthermore, by combining the function of recording object-related feedback and adjusting the priority of subsequent representation retrieval based on the feedback, the caching mechanism can be further optimized. For example, if a user provides negative feedback on a cached representation, the system can mark the cached item as "pending update" or "invalid," prompting the multi-level priority strategy to be re-executed on the next retrieval, and the priority adjusted based on the feedback, thereby ensuring the quality of cached content and user satisfaction. This mechanism enables the system to learn from user feedback and dynamically update and optimize cached content, forming an efficient and adaptive closed loop for object representation acquisition and management.

[0113] The following example illustrates this. Consider a smart assistant scenario where a user first says, "Show me a rose." The system first obtains the object information "rose." When attempting to retrieve its representation, the system first queries its internal cache. Since this is the first request, the image representation of "rose" may not exist in the cache. At this point, the system initiates a multi-level priority strategy, such as first attempting local matching, and if not found online, retrieving it. Ultimately, it may generate an image of a rose using an image generation model. Once the image is successfully generated and presented to the user, the system associates this rose image with the object information "rose" and stores it in the cache. When the user says "Show me a rose" again at a later time or in another session, the system again obtains the object information "rose" and queries the cache again. This time, since the image representation of "rose" already exists in the cache, the system can directly retrieve the image from the cache without performing time-consuming operations such as local matching, online retrieval, or content generation again, thus achieving a fast response.

[0114] Through the above technical solution, this application effectively solves the problems of resource waste, response latency, and low efficiency caused by the system repeatedly executing the retrieval process every time it encounters the same object. Specifically, when the system needs to obtain the representation of an object, it first queries the cache. If the representation of the object already exists in the cache, it can be directly retrieved and used, thereby avoiding the complex process of repeatedly executing multi-level priority strategies and significantly reducing the consumption of computing resources and network bandwidth. In addition, this caching mechanism greatly improves the system's response speed to repeated requests, providing users with a smoother and more immediate interactive experience. Combined with the multi-level priority strategy and feedback learning mechanism in the basic solution, the cache can store optimized and user-verified representations, further consolidating the system's learning results and ensuring that the subsequently obtained representations are of higher quality and better meet user expectations, thus constructing an efficient, intelligent, and adaptive object representation management system.

[0115] In some of the solutions described above in this application, a multi-level priority strategy is proposed to obtain object representations and record feedback to optimize subsequent acquisitions. However, in this process, the system only triggers representation generation when requested by the user, lacking the ability to proactively trigger it in specific scenarios. This results in the inability to proactively provide representations when the user may be interested or need them, thus limiting the intelligence and emotional connection of the interaction. To address this, this application further proposes to include proactively triggering object representation generation in specific scenarios and inquiring about user needs.

[0116] The phrase "in a specific scenario" as used in this application refers to the system identifying, based on preset rules or real-time analysis, the appropriate time or environment to proactively provide object representations to the user. This identification can be based on various factors, such as preset time points (e.g., the user's birthday, anniversary, or holiday), user habits (e.g., the user's first login after a long period of inactivity, or their operating patterns in a specific application), or by analyzing the user's current contextual information (e.g., geographical location, ongoing activities, or device sensor data). Alternatively, the system can utilize machine learning models to learn from user behavior and environmental data, identifying which object representations the user might be interested in, thereby intelligently determining the specific scenario.

[0117] "Proactively triggered object representation generation" refers to the system automatically initiating the object representation acquisition process upon recognizing a specific scenario, without explicit user instruction. Based on its internal logic or preset rules, the system autonomously calls the information acquisition module to collect information related to the predicted object and then activates the representation acquisition module, employing the multi-level priority strategy described in this application to generate the object's representation. This proactive triggering can be executed immediately upon recognizing a specific scenario, ensuring the system can quickly respond to potential user needs.

[0118] "Asking for user needs" refers to the process after the system actively triggers object representation generation. To ensure the generated representation meets the user's actual expectations and preferences, the system interacts with the user to obtain feedback or confirmation. This interaction can take various forms. For example, the system can show the user one or more generated representation options and ask, "Which one do you like?" or "Do you need any adjustments?". Alternatively, the system can pose an open-ended question to the user before presenting the representation, such as, "I've prepared a surprise for you regarding [object], what style would you like it to be?" or "Do you have any special requirements for the representation of [object]?" This incorporates the user's personalized needs during the generation phase. This approach effectively avoids generating content that does not meet user expectations and improves user satisfaction.

[0119] When the system identifies a specific scenario, such as detecting an approaching anniversary, a user restarting the system after a long period of inactivity, or determining based on user habits that the user might be interested in certain information, the system no longer passively waits but actively triggers object representation generation. This proactive behavior means that the system autonomously initiates the object representation acquisition process, employing the multi-level priority strategy described in this application, i.e., sequentially attempting local matching, online acquisition, or content generation to ensure efficient and successful acquisition of the object representation. After actively triggering and generating the representation, the system interacts with the user, for example through prompts, options, or dialogue, soliciting the user's opinions, preferences, or whether further adjustments are needed. This mechanism effectively bridges the system's proactivity with the user's personalized needs, ensuring that the system provides convenience while fully respecting the user's right to choose and customized requirements.

[0120] As a specific implementation, we can envision an intelligent assistant system. When the system identifies an upcoming wedding anniversary by analyzing a user's calendar or social media information, this constitutes a specific scenario. In this scenario, the system proactively initiates an object representation generation process. For example, based on the user's past preferences or anniversary theme, the system utilizes its multi-level priority strategy to first attempt to match a visual representation of a bouquet of roses in its local resource library. If not found, it searches online for an image. If that also fails, it calls an image generation model to generate an image of a bouquet of roses, or generates an audio representation of a blessing. After the representation is generated, the system can send a message to the user, such as, "I detected that your wedding anniversary is approaching. I've prepared a little surprise for you. Would you like to check it now?" This embodies the act of inquiring about the user's needs. If the user chooses to view it, the system presents the generated representation. The user's feedback on the representation (such as likes or suggestions for modification) is recorded and used to optimize future proactive triggering strategies and the accuracy of generated content.

[0121] Through the above technical solution, this application effectively addresses the problem that traditional systems cannot proactively respond in specific scenarios, resulting in limited interactive intelligence and emotional connection. The system intelligently identifies specific scenarios and proactively triggers object representation generation, enabling it to anticipate user needs and prepare relevant object representations before the user explicitly requests them. This significantly improves the system's response efficiency and the smoothness of the user experience. Simultaneously, the step of inquiring about user needs ensures that this proactivity is not blind or intrusive, but rather adjusted based on the user's immediate feedback and preferences, making the provided representations more accurate and personalized. This proactive and personalized interaction mode not only enhances the system's intelligence, enabling it to serve users more attentively, but also significantly improves the emotional connection between the user and the system by demonstrating foresight and care for user needs, making the interactive experience more natural and warm. Combining the multi-level priority acquisition and feedback learning mechanism of the basic method, this solution ensures that the proactively generated representations have a high success rate and high matching degree, and can be continuously optimized, thereby providing users with a more intelligent, personalized, and emotionally rich interactive experience.

[0122] Traditional object representation generation systems typically rely on a single source (such as a local icon library or database) to obtain object representations. If the representation is not found in that source, the system fails and cannot compensate through other means. Furthermore, the system lacks the ability to learn from user feedback on the representation results. The same acquisition process must be repeated every time the same object is encountered, failing to optimize the user experience and resulting in low interaction efficiency and insufficient personalization.

[0123] To address this issue, this application proposes an object representation generation system, including an information acquisition module, a representation acquisition module, and a feedback learning module. The information acquisition module acquires information related to the object; the representation acquisition module acquires the object's representation using a multi-level priority strategy based on the information; and the feedback learning module records object-related feedback and adjusts the priority of subsequent representation acquisitions based on the feedback. The core innovation of this embodiment lies in organically combining a multi-level priority strategy with a feedback learning mechanism. This allows for automatic switching to other methods when one acquisition method fails and dynamic optimization of the strategy based on user feedback, thereby improving the success rate of object representation acquisition and achieving personalized representation generation.

[0124] The information acquisition module is configured to receive information related to the target object. This information can come from user input (such as text or voice), other systems, smart agents, or devices. The information may include the object name, attribute description, and context. The role of the information acquisition module is to provide basic input for subsequent representation generation, ensuring that the system can initiate processing for a specific object and avoid blind operations. For example, when a user issues the command "play rain sound," the information acquisition module parses the target object as "rain sound" and its attribute description, providing a clear basis for representation acquisition.

[0125] The representation acquisition module is configured to acquire the representation of the object using a multi-level priority strategy based on the information provided. This multi-level priority strategy includes sequentially trying multiple representation acquisition methods: first, matching the object in a predefined local resource library; if it exists, returning it directly; if local matching fails, retrieving relevant resources via online search; if online acquisition fails or personalized content is required, dynamically generating the representation using a content generation model; if all methods fail, returning a general representation or a prompt message. This strategy improves the success rate and coverage by sequentially trying different acquisition methods, ensuring automatic switching when one method fails, thus overcoming the limitations of a single source. For example, when acquiring a "rose" image, the system prioritizes trying the local icon library, switches to online search after failure, and finally creates the image using a content generation model, forming a complete acquisition chain.

[0126] The feedback learning module is configured to record feedback related to the object (such as user likes, dislikes, and manual corrections) and dynamically adjust the priority of subsequent object retrieval based on this feedback. For example, if a user frequently dislikes an object, the weight of the corresponding retrieval method is reduced, and other methods are tried first. This implements a closed-loop learning mechanism, dynamically optimizing strategies from user interactions to improve the accuracy and personalization of subsequent responses. Specifically, when a user is dissatisfied with a "rose" image obtained through online search, the system records negative feedback and automatically reduces the priority of online search when retrieving the same object representation, instead prioritizing content generation methods, thereby gradually adapting to user preferences.

[0127] The following example illustrates this. In a smart home system, a user says "play the sound of ocean waves." The information acquisition module identifies the target object as "the sound of ocean waves" and its attributes; the representation acquisition module first matches it in the local audio library; if not found, it searches online for audio resources; if that also fails, it calls the audio generation model to create one; after the user criticizes the quality of the generated audio, the feedback learning module records this negative feedback and dynamically adjusts the priority strategy when acquiring "the sound of ocean waves" in the future, prioritizing the online search method. Through the above technical solution, the system can flexibly adopt multi-level strategies to acquire object representations, learn and optimize from user feedback, and support multi-modal representation generation, thereby improving the intelligence of interaction and user experience. Specifically, the multi-level priority strategy ensures that most objects can obtain representations without pre-setting all content; the feedback learning mechanism allows the system to continuously adapt to user preferences, becoming more user-savvy with use; the modular design enables the system to generate object representations efficiently and personalizedly, significantly improving interaction efficiency and user satisfaction.

[0128] In some of the embodiments described above in this application, a multi-level priority strategy is proposed to flexibly obtain object representations. However, in this process, the execution of the strategy may lack a clear attempt order, resulting in unreasonable resource allocation, reduced acquisition efficiency, or increased failure rate, and it cannot ensure that the higher priority method is executed first.

[0129] In this regard, this application further proposes that the representation acquisition module be configured to try multiple representation acquisition methods in sequence.

[0130] The representation acquisition module is the core component of the object representation generation system. Its main function is to acquire representations related to a specific object through various means based on object-related information provided by the information acquisition module. This module can be an independent software service, a processing unit, or a set of collaborative algorithms, responsible for coordinating the invocation of different acquisition methods and the processing of results. "Trying in sequence" means that when the representation acquisition module performs an acquisition task, it calls different representation acquisition methods one by one in a preset and explicit order. This means that only when the previous method fails or cannot meet the requirements will the next method be initiated, rather than simultaneously or randomly selected. This sequentiality ensures efficient use of resources and effective execution of the strategy. "Multiple representation acquisition methods" refers to different technical approaches or data sources used to acquire object representations. These methods can include, but are not limited to, matching and retrieving from a locally stored resource library, querying and acquiring from an external database or service via a network connection, or creating new content in real time using a generative model. For example, in addition to local matching, network acquisition, and content generation, it can also include selecting from a preset template library, generating through user-defined rules, or calling third-party application programming interfaces (APIs).

[0131] The proposed solution configures the representation acquisition module to sequentially attempt multiple representation acquisition methods, making the execution of a multi-level priority strategy in the object representation generation system orderly and efficient. Specifically, after the information acquisition module obtains object-related information, the representation acquisition module receives this information and, according to a preset priority order, first attempts the first representation acquisition method. If this method successfully acquires a satisfactory object representation, the attempt stops and the result is returned; if it fails, it automatically switches to the second representation acquisition method, and so on, until successful acquisition or all methods have been tried. This mechanism ensures that the system always prioritizes methods with lower costs, higher efficiency, or higher success rates, such as prioritizing matching from the local resource library to reduce network latency and computational resource consumption. Only when a higher-priority method fails to meet the requirements will the system advance to a lower-priority method that may be more flexible or have a wider coverage, such as network acquisition or content generation. This orderly trial process, combined with feedback recorded by the feedback learning module, can further optimize subsequent priority adjustments, thus forming an adaptive and highly efficient closed loop for object representation acquisition.

[0132] The following example illustrates this. In a smart assistant system, when a user issues the command "Show me a rose," the information acquisition module identifies the object as "rose." The acquisition module then initiates and attempts to acquire the visual representation of "rose" sequentially according to a preset priority order. First, the acquisition module attempts local matching, searching for an image of "rose" in the system's built-in local icon library or image library. If a local match is successful, the image is returned directly. If a local match fails, the acquisition module then attempts online acquisition, for example, by calling a search engine API or online image library interface to search for an image of "rose." If online acquisition is successful, the acquired image is returned. If online acquisition still fails, or the system determines that a more personalized image is needed (e.g., the user specifies a certain style), the acquisition module attempts content generation, for example, by calling an image generation model (such as a diffusion model or a generative adversarial network (GAN)) to generate an image of a "rose" in real time. Through this sequential trial mechanism, the system can ensure that the representation of the object is acquired in the most efficient and reliable way among various acquisition methods.

[0133] Through the above technical solution, this application effectively solves the problem of lacking a clear trial order during the execution of multi-level priority strategies. The representation acquisition module ensures that the multi-level priority strategy can be executed efficiently according to the preset logic by sequentially trying multiple representation acquisition methods, avoiding resource allocation chaos and unnecessary computational overhead. This orderly execution process significantly improves the success rate and efficiency of object representation acquisition, reduces system response latency, and provides a stable foundation for subsequent feedback learning and strategy optimization. Simultaneously, it enables the system to flexibly switch and compensate for the limitations of different acquisition methods, thereby improving system robustness and user experience.

[0134] In some of the embodiments described above in this application, a variety of representation acquisition methods are proposed to attempt to acquire object representations in turn. However, in this process, if the specific acquisition method is not clearly defined, the system may not be able to cover all possible sources, resulting in low acquisition efficiency or high failure rate, and failing to meet diverse object representation needs.

[0135] In this regard, this application further proposes several representation acquisition methods, including at least one of local matching, network acquisition, or content generation.

[0136] Specifically, local matching refers to the representation acquisition module retrieving object representations by searching a resource library pre-stored on the local device or within the system. This can be implemented by maintaining a database containing representations of commonly used icons, basic sound effects, standard haptic patterns, etc., in the device's built-in storage, allowing for quick searching and matching directly in this database when object information is received; alternatively, the system can utilize a local caching mechanism to store recently or frequently accessed object representations for subsequent quick retrieval. Network acquisition refers to the representation acquisition module accessing external resources via a network connection to obtain object representations when local matching fails. This can be implemented by sending query requests to cloud servers or third-party content providers via public or private online APIs to obtain representations such as images, audio, video, or 3D models related to the object; alternatively, the system can integrate a web search engine to search the internet based on the object description and filter out suitable representation resources. Content generation refers to the representation acquisition module dynamically creating object representations using algorithms or models when neither local matching nor network acquisition can meet the requirements. This approach can be implemented by using deep learning-based generative models (e.g., Generative Adversarial Networks (GANs) or Diffusion Models) to synthesize visual images, audio clips, or tactile signals based on text descriptions or attribute parameters; or by employing rule-based procedural generation techniques to dynamically construct object representations that meet specific requirements based on preset logic and input parameters.

[0137] Based on the aforementioned system architecture, after receiving object-related information from the information acquisition module, the representation acquisition module will initiate its configured multi-level priority strategy to obtain the object's representation. This strategy is not a single-path approach but rather ensures success and efficiency by sequentially trying multiple representation acquisition methods. Specifically, the representation acquisition module will first attempt to perform local matching, that is, search for a representation matching the current object in the system's internal or device-local resource library. If a local match is successful, the representation is directly adopted, thereby achieving rapid response and reducing dependence on external resources. However, when local resources cannot provide the required representation, the representation acquisition module will automatically switch to network acquisition, accessing external databases or online services via network connection to expand the search scope and obtain richer representation resources. If network acquisition still fails to obtain satisfactory results, or when personalized or customized representations need to be generated, the representation acquisition module will further initiate a content generation mechanism, using advanced generation algorithms or models to dynamically create entirely new representations based on the object's description and contextual information. This progressive and flexible switching strategy enables the system to effectively handle various complex scenarios, such as insufficient local resources, limited network resources, or the need for innovative content, greatly improving the success rate and adaptability of object representation acquisition. Simultaneously, the feedback information recorded by the feedback learning module can dynamically adjust the priority of these acquisition methods, further optimizing the efficiency of representation acquisition and user satisfaction.

[0138] As a specific implementation, suppose a user sends a command to the system through the information acquisition module: "Please show an ancient Chinese porcelain piece." Upon receiving this command, the acquisition module will activate its multi-level priority strategy. First, the acquisition module will attempt local matching, for example, searching a pre-stored local artifact image database for an image of "ancient Chinese porcelain." If no match is found in the local database, or the matched image is of poor quality, the acquisition module will automatically switch to online acquisition mode, attempting to retrieve relevant porcelain images from the internet by calling an online artifact image library API or performing a web image search. If the online acquisition results are still unsatisfactory—for example, failing to find a specific porcelain piece matching the characteristics of "ancient" and "Chinese," or if the user desires a unique, never-before-seen style of porcelain—the acquisition module will further activate the content generation function, using a pre-trained image generation model to dynamically generate a porcelain image that meets the requirements based on the description "ancient Chinese porcelain."

[0139] Through the above technical solution, this application effectively solves the problems of traditional systems in object representation acquisition, such as a single acquisition method, low coverage, high failure rate, and difficulty in adapting to diverse needs. Specifically, by introducing three diverse representation acquisition methods—local matching, network acquisition, and content generation—and combining them with a sequential trial mechanism of the representation acquisition module, the system can first utilize efficient local resources for rapid response, seamlessly expand to broad network resources when local resources are insufficient, and finally ensure the availability of representation through dynamic generation when existing resources are lacking or personalized content is required. This multi-level, multi-source acquisition strategy significantly improves the success rate and coverage of object representation acquisition, enabling the system to flexibly respond to various object representation needs, thereby providing users with a more intelligent, efficient, and personalized interactive experience.

[0140] In some of the solutions mentioned above in this application, content generation is proposed as a way to obtain object representation. However, in this process, the specific implementation of content generation may lack a clear modal support mechanism, which leads to the generated representation being limited to a single form and unable to cover multiple interaction needs such as visual, audio, tactile or action, thus making it difficult to meet the comprehensive representation requirements of diverse scenarios.

[0141] In this regard, this application further proposes content generation including, but not limited to, at least one of a visual content generation unit, an audio content generation unit, a tactile content generation unit, or a motion content generation unit.

[0142] Content generation refers to the dynamic creation of new object representations that meet specific requirements based on input information using algorithms or models. Its function is to flexibly generate content when the required representation is not available in a pre-defined resource library or when a highly personalized representation is needed. Content generation can be implemented using various technologies. For example, deep learning models such as Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), or Diffusion Models can be used to generate images, audio, or motion sequences; rule-based systems or parameterized models can also be used to generate tactile signals or simple graphics.

[0143] The visual content generation unit is a module specifically designed to generate visual representations of objects. This includes, but is not limited to, images, video clips, 3D models, or graphical icons. For example, this unit could be an image generation model, such as DALL-E based on the Transformer architecture or Stable Diffusion based on the diffusion model, capable of generating high-resolution images from text descriptions; or a 3D modeling engine capable of generating three-dimensional models from object attributes.

[0144] The audio content generation unit is a module specifically designed to generate auditory representations of objects. This includes, but is not limited to, speech, sound effects, background music, or ambient sounds. For example, this unit could be a text-to-speech (TTS) synthesizer that converts text into natural speech; or a sound effects synthesizer that generates ambient sound effects such as ocean waves or rain based on a description; or a music generation model that generates background music based on emotional or stylistic requirements.

[0145] The haptic content generation unit is a module specifically designed to generate haptic representations of objects. This includes, but is not limited to, vibration patterns, pressure feedback, or texture simulation. For example, this unit could be a haptic signal generation algorithm that converts the physical properties of an object (such as roughness or softness) into specific vibration frequencies and intensity patterns; or a machine learning-based model that generates corresponding haptic feedback signals based on visual input.

[0146] The motion content generation unit is a module specifically designed to generate object representations of dynamic behaviors or pose sequences. This includes, but is not limited to, animations, gestures, facial expressions, or motion trajectories of people or robots. For example, this unit could be a motion synthesizer capable of generating a coherent sequence of actions for a virtual character based on text instructions or intentions; or an inverse kinematics (IK) solver capable of calculating joint angles based on a target pose; or a module based on a large language model (LLM) that directly outputs motion parameters.

[0147] The proposed solution refines content generation into dedicated generation units for different modalities, allowing the representation acquisition module to flexibly invoke the appropriate generation unit based on specific needs when employing a multi-level priority strategy for content generation. When the system needs to generate visual representations, it can activate the visual content generation unit; when it needs audio representations, it can activate the audio content generation unit; when it needs haptic representations, it can activate the haptic content generation unit; and when it needs motion representations, it can activate the motion content generation unit. This design ensures that content generation is no longer single-modal but supports the generation of representations across multiple modalities, including visual, audio, haptic, and motion. In this way, the system can dynamically generate the most suitable multimodal object representation based on user input, context, or the needs of a specific scenario, thereby greatly improving the system's adaptability and expressiveness in complex interactive scenarios.

[0148] As a specific implementation, in a virtual reality (VR) shopping scenario, a user might say via voice command, "I want a red plush toy." At this point, the information acquisition module receives this command. This indicates that after failing to match locally or acquire data online, the acquisition module triggers content generation. To meet the user's visual and tactile perception needs for the "plush toy," the acquisition module simultaneously invokes both the visual content generation unit and the tactile content generation unit. The visual content generation unit can generate a 3D model or high-resolution image of a red plush toy and present it in the user's virtual field of vision. Simultaneously, the tactile content generation unit, based on the "plush" attribute, generates vibration patterns simulating a soft, fluffy feel and provides feedback through the user's VR controller or haptic gloves. Furthermore, if the toy is interactive, the motion content generation unit can generate a simple "waving" animation to make the toy appear more lifelike.

[0149] Through the aforementioned technical solutions, content generation is no longer limited to a single modality, but can flexibly support the generation of various forms of representation, such as visual, audio, tactile, or motion-based representations. This significantly enhances the comprehensiveness and flexibility of the system's generated object representations, enabling it to efficiently adapt to diverse interaction needs, thereby providing users with a richer, more immersive, and personalized interactive experience.

[0150] In some of the solutions described above in this application, a content generation unit is proposed to generate an object representation. However, in this process, the generated content may be inconsistent with the style of the target object or scene, resulting in a poor user experience and insufficient immersion.

[0151] In this regard, this application further proposes to include a style control module for applying preset style parameters during content generation.

[0152] The style control module is a functional unit whose core function is to manage and apply style constraints during the content generation process. This module can be a standalone software module integrated into the representation acquisition module, responsible for receiving output requests from the content generation unit and adjusting the input or internal parameters of the generation model according to preset style parameters before calling the content generation model. Alternatively, the style control module can be a hardware acceleration unit specifically designed to handle style transfer or style transformation algorithms, working in conjunction with the content generation unit to ensure stylistic consistency in the generated content. Content generation refers to the process of dynamically creating visual, audio, tactile, or motion representations through algorithms or models. This can be achieved by using deep learning models (such as generative adversarial networks or diffusion models) to generate images, audio waveforms, tactile vibration patterns, or motion sequences based on text descriptions or specific inputs. It can also be achieved through parametric models that synthesize corresponding multimedia content based on input parameters (such as color, texture, pitch, rhythm, vibration frequency, and joint angles). Preset style parameters are set values ​​or datasets used to guide the content generation process, ensuring that its output conforms to specific aesthetic, emotional, or functional requirements. These parameters can be user preference vectors, such as the "cartoon style" or "realistic style" selected by the user in the system settings, or timbre parameters applied in audio generation, or action styles (such as elegant or lively) applied in motion generation. These preferences are quantified as parameters. In addition, preset style parameters can also be scene templates. For example, in a "sci-fi scene," all generated content should have a metallic texture and cool color tone; in a "natural scene," it should favor organic forms and warm color tone. These templates contain a series of style feature definitions.

[0153] This application's solution addresses the issue of inconsistent content styles by introducing a style control module and applying preset style parameters during content generation. Specifically, when the representation acquisition module determines, based on a multi-level priority strategy, that an object's representation needs to be acquired through content generation, the corresponding visual content generation unit, audio content generation unit, haptic content generation unit, or motion content generation unit is activated. During this process, the style control module intervenes. It doesn't simply generate content but guides or constrains the generation process based on pre-defined style parameters. For example, if the system needs to generate a visual representation of an object, and the user or the current scene sets specific style requirements (such as "watercolor style" or "retro style"), the style control module will pass these style requirements as preset style parameters to the visual content generation unit. When generating images, the visual content generation unit adjusts its internal algorithm based on these parameters to ensure that the final output image has the specified watercolor or retro art style. This mechanism ensures that even dynamically generated content maintains a high degree of consistency with the user's expectations, the application scenario, or the inherent style of the target object, avoiding the problem of generated content being out of context. In this way, the solution proposed in this application can provide highly customized and consistent representations, thereby improving the overall success rate of representation acquisition and user satisfaction.

[0154] The following example illustrates this. In a virtual social scenario, a user wants an agent to display a "victory gesture," and the agent is currently set to an "elegant" style. After the system recognizes the "victory gesture," the acquisition module first attempts to match it in a predefined action library. If no victory gesture matching the "elegant" style is found, it may attempt an online search. If the online search also fails to provide results that meet the style requirements, the acquisition module will trigger the action content generation unit to generate content. At this point, the style control module receives the preset style parameter "elegant" and applies it to the action content generation unit. When generating the victory gesture animation, the action content generation unit adjusts the gesture's amplitude, speed, joint angles, etc., according to the "elegant" style parameter, making it exhibit smooth, soft, and rhythmic characteristics, rather than stiff or exaggerated movements. Ultimately, the victory gesture animation executed by the agent-driven Live2D character will perfectly match its "elegant" style setting, achieving a vivid and consistent interaction.

[0155] Through the aforementioned technical solutions, the style control module applies preset style parameters during content generation, effectively resolving the issue of inconsistent styles between generated content and the target object or scene. This significantly enhances the personalization and immersion of the user experience, ensuring that the generated visual, audio, tactile, or motion representations conform to specific style requirements. For example, in intelligent assistant dialogue, the generated object images maintain consistency with the assistant's overall style; in a virtual reality environment, the generated tactile feedback matches the material feel of virtual objects. This style consistency avoids interaction interruptions or discomfort caused by abrupt content style changes, allowing users to interact with the system more naturally and enjoyably. Combined with the multi-level priority strategy of the representation acquisition module, even when local and network resources are insufficient to provide the required style content, content generation can still provide high-quality, style-consistent representations through the style control module, thereby improving the overall success rate of representation acquisition and user satisfaction.

[0156] In some of the solutions mentioned above in this application, system modules are proposed for acquiring and adjusting object representations. However, when presenting object representations, there is a lack of ways to associate them with objects, resulting in stiff, unnatural, or out-of-context presentations that affect user experience and interactive immersion.

[0157] In this regard, this application further proposes to include a presentation module for presenting the representation in a manner associated with the object.

[0158] The presentation module is a functional unit in the system used to output the acquired object representation to the user or external environment. This module can be a software component responsible for calling the underlying operating system or hardware interfaces to control various output devices, such as displays, speakers, and haptic feedback devices. Alternatively, the presentation module can also include specific hardware interface circuits, such as display drivers, audio amplifiers, or haptic motor controllers. These hardware units work in conjunction with the software logic to complete the output of the representation. The core function of the presentation module is to serve as the outlet for information interaction between the system and the user, ensuring that the object representation can be effectively perceived by the user.

[0159] Presenting the representation in a manner associated with the object means that the method of presentation is not isolated or arbitrary, but rather customized and contextualized based on the inherent characteristics of the object, its context, or the user's specific intentions. Specifically, this can be manifested in intelligently selecting the most suitable output modality (such as visual, auditory, tactile, etc.) and the corresponding presentation medium (such as a screen, speaker, haptic feedback device) based on the type of object (e.g., a visual object like a "rose" or an auditory object like "the sound of waves") and the corresponding presentation medium. Furthermore, the details of the presentation can be finely adjusted according to the specific scene or contextual information of the object. For example, in interactive scenarios involving virtual characters, the object representation can be organically combined with the specific actions of the virtual character (such as "picking up" an object), making the presentation more natural. This associative presentation method aims to enhance the naturalness, immersion, and coherence of the user experience, effectively avoiding the abruptness or detachment from context in the presentation.

[0160] This application's solution introduces a presentation module, which presents the representation in a way that is associated with the object, forming a complete closed loop with the basic object representation generation system. First, the information acquisition module acquires initial information related to the object, providing context for subsequent representation generation and presentation. Next, based on this information, the representation acquisition module generates the object's representation using a multi-level priority strategy. Subsequently, the presentation module receives the generated representation and, combined with the original object information obtained from the information acquisition module or the context inferred by the system, intelligently selects and executes the most appropriate presentation method. For example, if the object information indicates that it is an item requiring visual presentation, and the current scene involves a virtual character, the presentation module will coordinate the virtual character's actions with the display of the visual representation. This closely integrated presentation method ensures that the generated representation is no longer an isolated fragment of information, but rather integrated with the object and its environment, greatly enhancing the realism and immersion of the interaction. Simultaneously, the feedback learning module can record user feedback on the presentation effect and adjust the subsequent representation generation strategy and presentation method accordingly, thereby achieving continuous optimization and enabling the system to better understand and meet the user's personalized needs.

[0161] The following example illustrates this. In a smart assistant application, when a user issues the command "Send me a rose," the system's information acquisition module identifies the object as a "rose." This acquisition module, based on a multi-level priority strategy, may obtain an image representation of the rose through local matching, online acquisition, or content generation. The presentation module then receives this image representation and, considering the characteristics of the "rose" object and the context of the smart assistant as a virtual character, intelligently selects the presentation method. Specifically, the presentation module can drive the smart assistant's Live2D character to perform a "holding up with both hands" action, while simultaneously displaying the generated rose image precisely at the virtual character's hand position. The user appreciates this vivid and natural presentation, and the feedback learning module records this positive feedback. Consequently, when encountering similar items in the future, the system will prioritize this visual-action-linked presentation strategy.

[0162] Through the above technical solutions, the system can present object representations to users in a more natural and contextualized manner, effectively solving the problems of stiff, unnatural, or out-of-context presentations in traditional systems. This associative presentation not only enhances the fluency and immersion of the user experience but also enables users to perceive and understand objects more intuitively and realistically, thereby significantly enhancing the intelligence and effectiveness of human-computer interaction.

[0163] In some of the solutions mentioned above in this application, a representation acquisition module is proposed to acquire the representation of an object using a multi-level priority strategy based on information. However, in this process, the system needs to repeatedly execute the acquisition process every time it encounters the same object, resulting in wasted resources, response delays, and may increase network load, affecting user experience.

[0164] In this regard, this application further proposes that the object representation generation system also includes a caching module for storing the acquired representations. The caching module is a component specifically designed for temporary or persistent data storage. Its core concept is to save successfully acquired object representations so that they can be quickly retrieved and reused when needed later, thereby avoiding repeated, time-consuming acquisition processes. As one possible implementation, the caching module can be a high-speed memory, such as random access memory (RAM), to store frequently accessed representations for extremely fast read and write speeds. As another possible implementation, the caching module can also be a persistent storage medium, such as a solid-state drive (SSD) or non-volatile memory (NVM), to permanently store representations, ensuring data retention even after a system restart. Furthermore, in distributed systems, the caching module can also employ a distributed caching system, such as Redis or Memcached, to support large-scale concurrent access and high availability.

[0165] This caching module works closely with the information acquisition module, representation acquisition module, and feedback learning module in the system. Specifically, after the information acquisition module obtains information related to an object, the representation acquisition module uses a multi-level priority strategy to obtain the representation of that object based on the information. Once the representation acquisition module successfully obtains the representation of an object, the representation is stored by the caching module and associated with the object's related information. When the information acquisition module obtains the same or similar object information again, the representation acquisition module will first query the caching module before initiating its multi-level priority acquisition process (including local matching, network acquisition, or content generation). If the caching module already has a stored representation corresponding to the object information, the representation acquisition module can directly read and return the representation from the cache, thus skipping the time-consuming multi-level priority acquisition process. This mechanism significantly reduces the consumption of redundant computing resources and the number of network requests. Simultaneously, the feedback recorded by the feedback learning module can further guide the caching module's strategy. For example, for representations with poor user reviews, they can be removed from the cache or their priority reduced to ensure the quality of representations in the cache. In this way, the caching module, together with the information acquisition module, the representation acquisition module, and the feedback learning module, forms an efficient and adaptive closed loop for object representation acquisition, which greatly optimizes the overall system efficiency and user experience.

[0166] As a specific implementation, the aforementioned caching module can be a key-value pair-based in-memory database, such as Redis. When the representation retrieval module successfully retrieves a representation of an object—for example, if a user inputs "rose" through the information retrieval module, the module generates an image of a rose using content generation. This image data (or its storage path) is stored in the Redis cache as the value, with "rose" as the key. Simultaneously, metadata such as retrieval time, source, and user satisfaction ratings related to the feedback learning module can be associated. When the user inputs "rose" again through the information retrieval module, the representation retrieval module first queries Redis to see if a representation corresponding to "rose" exists. If it does, the image is retrieved directly from Redis and displayed, avoiding time-consuming operations such as local matching, online retrieval, or content generation. The cache invalidation strategy can be set to time-based, such as automatic invalidation after 24 hours, or feedback-based, such as immediate invalidation or degradation after a user downvotes a cached representation, to ensure the real-time nature and accuracy of the cached content.

[0167] Through the above technical solutions, the system effectively solves the problems of resource waste, response latency, and network burden caused by repeatedly executing the retrieval process when encountering the same object. The introduction of the caching module enables the efficient storage and reuse of already retrieved representations, thereby significantly reducing redundant calculations and network requests by the representation retrieval module and greatly improving system response speed and overall operating efficiency. This not only optimizes the user experience, allowing users to feel smoother and more immediate feedback when interacting with the system, but also reduces the system's consumption of computing resources and network bandwidth, achieving resource-friendly operation.

[0168] In some of the solutions described above in this application, an object representation generation system is proposed to generate object representations based on acquired information. However, in this process, the system only passively responds to requests and cannot proactively trigger object representation generation in specific scenarios. Therefore, it cannot provide relevant representations when the user may not have actively requested them, limiting the intelligence and emotional connection of the interaction. To address this, this application further proposes that the object representation generation system also include an active triggering module, used to proactively trigger object representation generation in specific scenarios and inquire about user needs.

[0169] The active triggering module is a functional component of the system whose main responsibility is to identify and respond to specific conditions or events, thereby initiating the object representation generation process. This module can be an independent software service or process that continuously monitors the system's internal state, user behavior patterns, or changes in the external environment; alternatively, it can be activated as part of the core control logic through a preset rule engine, scheduled tasks, or event listening mechanisms. Its core function is to transform the system from a passive response mode to an active service mode.

[0170] The aforementioned proactive triggering of object representation generation in specific scenarios refers to the system's ability to intelligently determine when to provide object representations to the user based on preset or dynamically identified contextual information. This identification of "specific scenarios" can be achieved in various ways. For example, it can be based on time events, such as preset dates like the user's birthday, anniversaries, or holidays; it can also be based on user behavior patterns, such as the system analyzing the user's usage habits in specific applications, their first login after a long period of inactivity, or identifying the user's current idle time; furthermore, it can be based on environmental context, such as a smart home system detecting that a user has returned home, or a smart vehicle system detecting that a user has arrived at a specific location; it can even combine external data sources, such as weather forecasts, news headlines, or social media trends, to determine the triggering timing.

[0171] The aforementioned inquiry into user needs refers to the system's proactive interaction with the user after actively triggering the generation of an object representation, in order to ensure that the provided representation matches the user's actual wishes and preferences, in order to obtain confirmation or further input. This inquiry can take various forms. For example, the system can ask the user questions through a dialogue interface, whether via text message or voice interaction, such as "We've detected that your wedding anniversary is approaching; would you like us to generate an image of a bouquet of roses for you?"; or, the system can use a graphical user interface, such as a pop-up prompt or notification on the screen, providing clear options for the user to choose from, such as "The system has prepared holiday greetings for you; would you like to view them?"; in more complex interaction scenarios, multimodal interaction methods can also be used, such as a virtual assistant simultaneously inquiring about the user through gestures and voice.

[0172] This application's solution introduces an active triggering module, enabling the object representation generation system to shift from a passive response to an active service. Specifically, when the active triggering module identifies a specific scenario, it intelligently determines and generates object information related to potential user needs based on the context of that scenario. For example, when it detects that a user's anniversary is approaching, the active triggering module can generate object information such as "roses" or "blessings." Subsequently, the system, based on this information and the results of user request queries, passes the determined object information to the information acquisition module. Upon receiving this information, the information acquisition module activates the representation acquisition module, which, according to its multi-level priority strategy, sequentially attempts local matching, online acquisition, or content generation to generate the corresponding representation for the object. During this process, the feedback learning module continues to record user feedback on these actively generated representations and adjusts the priority of subsequent representation acquisition accordingly, thus forming a closed-loop intelligent learning and optimization process. This mechanism allows the system not only to respond to explicit user instructions but also to provide timely and personalized services based on predictions of user intent and context, even without a user's active request. In this way, the active triggering module works closely with the information acquisition module, representation acquisition module, and feedback learning module to jointly build a more forward-looking and intelligent object representation generation system, which greatly improves the system's interactivity and user experience.

[0173] The following is a concrete example. As a specific implementation, in a smart assistant system, the system can be configured with an active trigger module to proactively trigger object representation generation in specific scenarios. For example, the system predicts that a user's wedding anniversary is approaching by analyzing the user's calendar information or social media data, thus constituting a "specific scenario." At this time, the active trigger module will proactively initiate the object representation generation process during the user's idle time, such as at night when the system load is low. It might first generate an image representation related to "roses" and an audio representation of a "blessing voice," and cache them. When generating these representations, the system can apply preset style parameters to ensure that the generated content is consistent with the user's preferences or the overall system style. Subsequently, on the anniversary, when the user interacts with the smart assistant, the system will present the pre-generated gift to the user through the active trigger module, such as displaying a rose image on the screen and playing a blessing voice, and may ask the user, "This is a surprise for your anniversary, do you like it?" If the user "likes" the surprise, the feedback learning module will record this positive feedback, thus prioritizing this proactive generation and presentation method when encountering similar scenarios in the future.

[0174] Through the aforementioned technical solutions, the system can transform from a traditional passive response mode to a proactive service mode, significantly enhancing the intelligence of the interaction and the user experience. The system no longer simply waits for explicit user commands, but proactively predicts and fulfills potential user needs in specific scenarios, such as anniversaries or when the user is at free time, providing relevant responses even when the user may not have actively requested them. This proactivity not only strengthens the emotional connection between the system and the user, making the interaction more human and surprising, but also, through its combination with a multi-level priority representation acquisition and feedback learning mechanism, ensures the quality and personalization of proactively generated content, further optimizing the user experience.

[0175] In some of the solutions mentioned above in this application, an object representation generation method is proposed to obtain object representations through a multi-level priority strategy and optimize based on feedback. However, in this process, the method itself lacks the support of persistent storage and executable carrier, which makes it impossible to be permanently saved, deployed across devices, or reused efficiently, thereby limiting the actual application scope and execution efficiency of the method.

[0176] In this regard, this application further proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements an object representation generation method.

[0177] The computer-readable storage medium refers to a physical or logical carrier capable of storing digital data or computer instructions. This medium can be a non-transitory storage medium, such as a hard disk drive (HDD), solid-state drive (SSD), read-only memory (ROM), random access memory (RAM), flash memory, optical disc (CD-ROM, DVD, Blu-ray disc), or magnetic tape, used to ensure that computer programs are preserved after power loss and to support program loading and execution. Alternatively, the medium can be a temporary storage medium, such as a transmission medium like electrical signals, optical signals, or electromagnetic waves, used to temporarily store programs during data transmission. The computer program refers to a set of instructions that, when executed by a computer, can complete a specific task or achieve a specific function. The program can exist in the form of source code, such as a text file written in high-level programming languages ​​like Python, Java, or C++, or in the form of compiled executable code, such as binary files, bytecode, or machine code. Its function is to transform abstract object representation methods into a sequence of instructions that the computer can understand and execute. The processor refers to the hardware unit that executes computer program instructions and is a core component of the computer system. The processor can be a central processing unit (CPU), such as the Intel Core series or AMD Ryzen series, or a graphics processing unit (GPU), such as the NVIDIA GeForce series or AMD Radeon series, especially suitable for parallel computing tasks. Alternatively, the processor can be an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a microcontroller (MCU), whose main responsibility is to parse and execute instructions in the computer program, driving the various steps of the object representation generation method. Implementing the object representation generation method means that the processor completes all steps and functions of the object representation generation method by executing the computer program. Specifically, the processor can sequentially execute steps such as acquiring information related to the object, acquiring the object's representation using a multi-level priority strategy, recording feedback, and adjusting priorities according to program instructions. Alternatively, the processor can call functions or modules encapsulated in the program to implement the various functional units of the object representation generation method in a modular way, thereby transforming the static program code on the storage medium into dynamic, actually running business logic.

[0178] This application's solution transforms abstract method logic into a deployable and executable entity by embedding the object representation generation method within a computer program stored on a computer-readable storage medium and executing it via a processor. The computer-readable storage medium provides persistent storage space for the computer program, ensuring the long-term preservation of the object representation generation method code and avoiding redundant operations that require reloading on each execution. The computer program stored on it transforms the complex logic of the object representation generation method into a series of instructions that can be recognized and executed by the processor. When the processor receives the execution instructions, it efficiently executes each step of the object representation generation method according to the preset logical order in the computer program. This includes acquiring object-related information, obtaining object representations according to a multi-level priority strategy, recording feedback, and adjusting the priority of subsequent representation acquisition based on the feedback. This collaborative working mechanism transforms the object representation generation method from a mere theoretical concept into a practical solution capable of stable and automated operation in various computing environments. In this way, the solution proposed in this application not only solves the problem of the lack of persistent storage and executable carrier in the method itself, but more importantly, it enables the multi-level priority strategy, feedback learning mechanism and multi-modal representation generation capability of the above object representation generation method to be implemented and optimized efficiently and reliably in practical applications, thereby significantly improving the practicality and intelligence level of the entire object representation generation scheme.

[0179] The following is a concrete example illustrating a specific implementation. Imagine a smart voice assistant system running on a smartphone. The computer-readable storage medium can be the smartphone's internal flash memory (e.g., eMMC or UFS storage), responsible for persistently storing all system files and application data. The computer program can be the compiled code of the smart voice assistant application (e.g., an APK file for Android or an IPA file for iOS), installed and stored on the flash memory, encapsulating the complete logic of the object representation generation method. When a user interacts with the smart assistant via voice commands (e.g., "Show me a rose"), the smartphone's system-on-a-chip (SoC), with its integrated central processing unit (CPU) and graphics processing unit (GPU), constitutes the processor. This processor executes the relevant code in the smart voice assistant application to implement the object representation generation method. Specifically, the processor first executes program instructions to obtain the user-inputted "rose" information, and then, according to a multi-level priority strategy defined in the program, sequentially attempts to obtain a visual representation of the rose from local cache, an online image library, or a generative AI model. After obtaining the representation and presenting it to the user, if the user performs a "like" or "dislike" action on the result, the processor will execute the feedback recording logic in the program, store the user feedback in the database on flash memory, and adjust the priority of obtaining similar object representations in the future based on the feedback.

[0180] Through the above technical solution, the object representation generation method is transformed from an abstract logical concept into a deployable and executable software entity, thus solving the problem of the lack of persistent storage and executable carrier for the method in practical applications. This solution ensures that the code of the object representation generation method can be stored for a long time, avoiding redundancy that requires reloading every time it is executed, and greatly improving the deployment efficiency and reusability of the method. In addition, through the efficient execution of the processor, each step of the object representation generation method can be completed automatically and quickly, significantly improving the system's response speed and overall performance. This implementation method enables the core functions of the above object representation generation method, such as multi-level priority strategy, feedback learning mechanism, and multimodal representation generation, to run stably and reliably, thereby giving full play to its advantages in practical applications and providing users with a more intelligent, personalized, and efficient interactive experience.

[0181] Other application scenarios The following simplified embodiments illustrate the application of the present invention in other fields. Specific implementations of these embodiments can be found in the examples described in the detailed embodiments above, and will not be repeated here.

[0182] Simplified Example 1: Multimodal Fusion (Virtual Shopping) The user says, "I want a red plush toy." The system simultaneously generates a 3D image of the toy (visual) and a soft, tactile feedback description (haptic), and optionally plays a soothing background music (audio). If the user is satisfied with the image but not the tactile feedback, they can tap the screen to adjust the tactile generation parameters.

[0183] Simplified Example 2: Generation of Robot Motion Representations In smart home or industrial scenarios, when a user says, "The robotic arm picks up the cup," the system recognizes the action as "picking up the cup," first checks a predefined action library, and if not found, searches online for similar action videos. If that also fails, it uses a motion generation model to generate the robotic arm's trajectory. The system adjusts the trajectory based on the current environment (the cup's position) and caches the action. The robot then executes the action, and user feedback is used for optimization.

[0184] Simplified Example 3: Multimodal Fusion Representation (Action + Haptic) In the VR experience, the user says, "Hold a soft ball." The system generates a visual representation (a 3D model of the ball), a tactile representation (a soft vibration pattern), and a motion representation (an animation of the hand holding the ball). These three are kept consistent through style control (e.g., the ball has a plush texture, and the holding motion is gentle). If the user is not satisfied with the holding motion, they can tap the screen, and the system will adjust the motion generation parameters.

[0185] Simplified Example 4: Generation of Industrial Equipment Status Icons When the operator asks, "What's the machine temperature?" the system identifies the object as "temperature," retrieves the temperature value from the equipment data, generates a thermometer icon from the icon library, and displays it on the control panel. The system records the operator's taps, and a more intuitive dashboard display will be used later.

[0186] The above descriptions are merely embodiments of this application and are not intended to limit the scope of protection of this application. It should be noted that the term "user" in this application is a preferred embodiment of the interactive object, but not a limitation on the scope of protection. Those skilled in the art should understand that the interactive object can be any entity capable of interacting with the interactive system, such as a human user, other systems, devices, or applications. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

[0187] While this solution is applicable to local devices, it can also be deployed on cloud servers. In a cloud environment, larger-scale embedded models and vector databases can be used to further improve retrieval accuracy and capacity. Those skilled in the art should understand that the technical solution of this invention is not limited to a specific deployment environment, and any implementation based on the technical concept of this application should be considered to fall within the protection scope of this invention.

Claims

1. A method of object representation generation, characterized by, Includes the following steps: Retrieve information related to the object; Based on the information, a multi-level priority strategy is used to obtain the representation of the object; Record feedback related to the object and adjust the priority of obtaining the representation in subsequent transactions based on the feedback.

2. The method of claim 1, wherein, The multi-level priority strategy includes trying multiple representation acquisition methods in sequence.

3. The method of claim 2, wherein, The various representation acquisition methods include at least one of local matching, online acquisition, or content generation.

4. The method of claim 3, wherein, The content generation includes, but is not limited to, at least one of visual content generation, audio content generation, haptic content generation, or motion content generation.

5. The method of claim 4, wherein, The content is generated using preset style parameters to ensure that the generated content matches the style of the object.

6. The method of claim 1, wherein, The representation includes, but is not limited to, at least one of visual representation, audio representation, tactile representation, or motion representation.

7. The method of claim 1, wherein, It also includes presenting the representation in a manner associated with the object.

8. The method of claim 1, wherein, The recorded feedback includes recording user feedback.

9. The method of claim 8, wherein, The user feedback includes, but is not limited to, at least one of the following: likes, dislikes, or manual corrections.

10. The method of claim 1, wherein, It also includes caching the information associated with the obtained representation and the object to avoid repeated retrieval.

11. The method of claim 1, wherein, It also includes proactively triggering object representation generation in specific scenarios and soliciting user requests.

12. An object representation generation system, characterized by, include: The information acquisition module is used to acquire information related to the object; The representation acquisition module is used to acquire the representation of the object based on the information and using a multi-level priority strategy; The feedback learning module is used to record feedback related to the object and adjust the priority of obtaining the representation in subsequent acquisitions based on the feedback.

13. The system of claim 12, wherein, The representation acquisition module is configured to try multiple representation acquisition methods in sequence.

14. The system of claim 13, wherein, The various representation acquisition methods include at least one of local matching, online acquisition, or content generation.

15. The system of claim 14, wherein, The content generation includes, but is not limited to, at least one of a visual content generation unit, an audio content generation unit, a tactile content generation unit, or a motion content generation unit.

16. The system of claim 15, wherein, It also includes a style control module, which applies preset style parameters when content is generated.

17. The system of claim 12, wherein, It also includes a presentation module for presenting the representation in a manner associated with the object.

18. The system of claim 12, wherein, It also includes a caching module for storing the retrieved representations.

19. The system of claim 12, wherein, It also includes an active triggering module, which is used to actively trigger the generation of object representations in specific scenarios.

20. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of any one of claims 1 to 11.