Body model representation method and related device
Patent Information
- Application Number
- CN202510901741.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2045-06-30
AI Technical Summary
然而,具身模型是具备复合能力的综合性模型,由多种类型的模型构成,利用相关技术的中间表示方案无法对其进行准确的中间表示
Smart Images

Figure CN120874974B_ABST
Abstract
Description
Technical Field
[0001] The embodiments described in this application relate to the field of computer technology, and in particular to embodied model representation methods and related devices. Background Technology
[0002] To meet user needs for model transfer, optimization, and visualization, intermediate representation schemes exist in related technologies. These intermediate representations can represent the model's structure, computational logic, and data flow using standardized and unified formats. Different models possess different characteristics, and related technologies typically employ structures adapted to these characteristics for accurate intermediate representation. However, embodied models are complex models with multiple capabilities, composed of various model types, and the intermediate representation schemes used in related technologies cannot accurately represent them. Summary of the Invention
[0003] In view of this, multiple embodiments of this application aim to provide a method and related equipment for representing embodied models, which can accurately represent embodied models and improve the accuracy of intermediate representation of embodied models.
[0004] One embodiment of this application provides an embodied model representation method, wherein the embodied model includes multiple sub-models corresponding to at least two model types; the method includes: determining intermediate representation information for each sub-model based on the intermediate representation information type corresponding to the model type of each sub-model; the intermediate representation information types corresponding to different model types are different; and determining intermediate representation information between different sub-models based on the computational relationship between different sub-models.
[0005] Optionally, the sub-model includes one or more computation nodes, and the intermediate representation information of the sub-model includes a computation graph; the step of determining the intermediate representation information of the sub-model includes: recording the input and output information of the computation nodes used by the sub-model to construct a computation graph corresponding to the sub-model.
[0006] Optionally, the step of recording the input and output information of the computing nodes used by the sub-model includes: recording the input information of the computing node using a first tracking function before inputting the input information into the computing node's computing function; and recording the output information of the computing node using a second tracking function after the computing node's computing function has calculated the output information.
[0007] Optionally, the method further includes: applying the target encapsulation standard to the computation graph construction process of each sub-model; wherein the target encapsulation standard is used to define the recording format of the input and output information of the computation nodes.
[0008] Optionally, the step of determining the intermediate representation information of each sub-model based on the intermediate representation information type corresponding to the model type of each sub-model includes: determining the intermediate representation information of each sub-model according to the segmentation marker in the embodied model and the intermediate representation information type corresponding to the model type of each sub-model; wherein the segmentation marker is used to segment multiple sub-models in the embodied model and / or multiple computation nodes in the sub-model.
[0009] Optionally, the splitting marker is added before or after the specified computation node of each sub-model in the embodied model; or, the splitting marker is added before or after the specified computation node of the splitting instruction.
[0010] Optionally, each sub-model includes a visual model and a large language model; the intermediate representation information of the visual model includes a computation graph in ONNX format or tvm.relax format; the intermediate representation information of the large language model includes a JSON file and model parameters; the intermediate representation information between different sub-models includes a computation graph in ONNX format, tvm.relax format, or prototext format.
[0011] One embodiment of this application also provides an embodied model representation device, wherein the embodied model includes multiple sub-models corresponding to at least two model types; the device includes: a first intermediate representation module, used to determine the intermediate representation information of each sub-model based on the intermediate representation information type corresponding to the model type of each sub-model; the intermediate representation information types corresponding to different model types are different; and a second intermediate representation module, used to determine the intermediate representation information between different sub-models based on the calculation relationship between different sub-models.
[0012] One embodiment of this application also provides a computer device, the computer device including a memory and a processor, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to implement the method as described above.
[0013] One embodiment of this application also provides a computer-readable storage medium storing at least one computer program that, when executed by a processor, can implement the method described above.
[0014] One embodiment of this application also provides a computer program product for implementing the method as described above.
[0015] In the various embodiments provided in this application, intermediate representation information for each sub-model is generated for the intermediate representation information type corresponding to the model type of each sub-model, and intermediate representation information between different sub-models is determined based on the calculation relationship between different sub-models. This can accurately realize the intermediate representation of the embodied model, fill the gap in the intermediate representation scheme of the embodied model, and improve the accuracy of the intermediate representation of the embodied model. Attached Figure Description
[0016] Figure 1 A system architecture diagram for an embodiment model representation method provided in one embodiment of this application.
[0017] Figure 2 A flowchart of an embodiment model representation method provided for one embodiment of this application.
[0018] Figure 3 A schematic diagram showing an intermediate representation of an embodied model provided for one embodiment of this application.
[0019] Figure 4 A computational diagram of a sub-model provided for one embodiment of this application.
[0020] Figure 5 This is a schematic diagram illustrating the call sequence of the tracing function provided in one embodiment of this application.
[0021] Figure 6 A schematic diagram of the segmentation markers provided for one embodiment of this application.
[0022] Figure 7 A structural diagram of an embodied model representation device provided for one embodiment of this application.
[0023] Figure 8 A schematic diagram of a computer device provided for one embodiment of this application. Detailed Implementation
[0024] The information to be retrieved in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0025] In the description of the embodiments of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0026] An embodied model is an artificial intelligence model that operates within an intelligent agent equipped with sensors and actuators, possessing capabilities such as perception, cognition, planning, and execution. This intelligent agent includes, but is not limited to, robots, virtual humans, and autonomous driving systems. Compared to AI models with a single capability (such as perception), embodied models can be used to realize a closed-loop action of the intelligent agent consisting of perception, decision-making, and action, thus having a wider range of applications.
[0027] For example, embodied models can be applied to complex scenarios such as intelligent robot control, virtual human interaction, AR / VR intelligent agents, autonomous driving, and simulation training. In these scenarios, it is usually necessary to process various heterogeneous data such as images, text, voice, and decision-making. Applying embodied models in these scenarios can achieve the expected motion goals (e.g., walking into the kitchen to grab a water glass).
[0028] Generally, embodied models include, but are not limited to, the following sub-models: visual models, large language models, control models, and multimodal fusion models. Visual models process data collected by the visual module, specifically performing tasks such as image recognition, object detection, and semantic segmentation. Large language models perform language interaction tasks such as understanding text, generating instructions, and planning dialogues. Control models generate control policies to drive agents to perform target behaviors. Multimodal fusion models integrate one or more multimodal data sources, such as images, text, speech, and actions, to generate corresponding decisions based on comprehensive data integration results.
[0029] Intermediate Representation (IR) refers to a standardized modeling approach that sits between the model's source code and the underlying execution environment. It can be used to represent the model's structure, computational logic, and data flow. Generally, IR is used to bridge different frameworks, platforms, and languages. Due to structural differences between different models, IR can be of various types (e.g., computation graph, control flow) to adapt to different types of models and achieve accurate model representation.
[0030] Because model training is performed on high-performance servers, but actual deployment may occur on edge devices, embedded systems, or mobile phones, the original training code generally cannot be run directly between different devices and inference engines. Therefore, related technologies use Inference Representation (IR) to standardize the model representation, enabling model transfer across different platforms. Furthermore, IR facilitates efficient model visualization, structural analysis, and structural optimization (e.g., operator fusion, constant folding, layer rearrangement).
[0031] When representing models with a single capability, related technologies can select appropriate intermediate representation information types based on the model's characteristics. However, embodied models are composed of diverse sub-models, each with different characteristics, thus requiring different types of intermediate representation information. For example, for visual models, an ONNX format computational graph is used for IR (Indirect Representation), while for large language models, a JSON file combined with model parameters is used. The IR solutions provided by related technologies are difficult to apply to embodied models.
[0032] Therefore, it is necessary to provide an embodied model representation method that generates intermediate representation information for each sub-model based on the intermediate representation information type corresponding to the model type of each sub-model, and determines the intermediate representation information between different sub-models based on the calculation relationship between different sub-models. This can accurately realize the intermediate representation of the embodied model, fill the gap in the intermediate representation scheme of the embodied model, and improve the accuracy of the intermediate representation of the embodied model.
[0033] Please see Figure 1 In several embodiments provided in this application, the embodied model representation method can be applied to an embodied model representation device.
[0034] In this embodiment, the embodied model representation device can be an electronic device with certain computing power and network access capabilities. This electronic device can be a desktop computer, laptop computer, tablet computer, or even a server. The electronic device can connect to the server via a network. The server can be a distributed server, including multiple processors, memory, network communication modules, etc., working together to achieve various functions. Alternatively, the server can also be a server cluster formed by several servers, possessing higher computing and data processing capabilities. With the development of science and technology, the server can also be implemented using new technological means, such as a new type of "server" based on quantum computing. Of course, in some embodiments, the embodied model representation device can also be a program module running in an electronic device.
[0035] In this embodiment, the electronic device may include, but is not limited to, a processor, a memory, and an intermediate representation information generation module. The processor is used to identify the model type of the sub-model and determine the intermediate representation information type corresponding to that model type. The memory is used to store the original representation information of each sub-model in the embodied model, the intermediate representation information of each sub-model, and the intermediate representation information between different sub-models. The intermediate representation information generation module is used to determine the intermediate representation information of each sub-model based on the intermediate representation information type corresponding to the model type of each sub-model, and to determine the intermediate representation information between different sub-models based on the computational relationships between them.
[0036] Please see Figure 2One embodiment of this application provides an embodied model representation method, wherein the embodied model includes multiple sub-models corresponding to at least two model types. The embodied model representation method can be applied to an embodied model representation device. The embodied model representation method may include the following steps.
[0037] Step S110: Determine the intermediate representation information of each sub-model based on the intermediate representation information type corresponding to the model type of each sub-model; the intermediate representation information types corresponding to different model types are different.
[0038] Step S120: Based on the computational relationship between different sub-models, determine the intermediate representation information between different sub-models.
[0039] In this embodiment, the embodied model has a composite model structure composed of multiple heterogeneous sub-models. In order to perform an accurate intermediate representation (IR) of the embodied model, each sub-model can be identified from the model construction information of the embodied model; or, the running embodied model can be tracked to identify each sub-model.
[0040] In this embodiment, the intermediate representation information of each sub-model can be generated according to the intermediate representation information type corresponding to the model type of each sub-model, and the intermediate representation information between different sub-models can be determined based on the calculation relationship between different sub-models.
[0041] Each sub-model is used to implement different functions; optionally, each sub-model can be used to implement perception, understanding, decision-making, and execution functions, respectively. The model types of sub-models include, but are not limited to: visual models, large language models, control models, and multimodal fusion models. Different model types correspond to different structural characteristics; therefore, the intermediate representation information types corresponding to different model types can be different. The intermediate representation information type indicates the representation language / format / method used to represent the corresponding model type. For example, intermediate representation information types include: onnx, tvm.relax, configuration files, and weight files.
[0042] In this embodiment, the intermediate representation information of the sub-model, as a specific representation result of the intermediate representation information type, can represent the sub-model in a standardized way. The intermediate representation information includes, but is not limited to, computational graphs, control flow, domain-specific languages, and weight combination configuration information. Optionally, when the sub-model is a feedforward neural network with a fixed structure that is easy to graphically visualize, a computational graph with a simple structure consisting of nodes and edges can be used as the intermediate representation information. Optionally, when the sub-model is a reinforcement learning model, control flow supporting if-else, loop, and branch structures can be used as the intermediate representation information. Optionally, when the sub-model is a deep learning compiler, a domain-specific language (e.g., TVM Script, MLIR) can be used as the intermediate representation information. Optionally, when the sub-model is a Huggingface large language model, BERT, or GPT, weight combination configuration information with a JSON / YAML defined structure and binary storage of weights can be used as the intermediate representation information.
[0043] In this embodiment, in order to comprehensively represent the embodied model, intermediate representation information between different sub-models is determined based on the computational relationships between them. This intermediate representation information represents the data dependencies, input / output paths, and upstream / downstream call relationships between the sub-models.
[0044] In some implementations, to comprehensively represent the intermediate representation results of the embodied model, the intermediate representation information of each sub-model, as well as the intermediate representation information between different sub-models, can be fused. The fused result is then displayed; optionally, this fused result can be represented as a computation graph, as detailed in [reference needed]. Figure 3 In this computation graph, the intermediate representation information of the sub-model is represented by an independent encapsulation unit (Scope), and the computational operation / function call inside the sub-model is represented by (Node).
[0045] In summary, this implementation generates intermediate representation information for each sub-model based on its model type and intermediate representation information type. Furthermore, it determines the intermediate representation information between different sub-models based on their computational relationships. This accurately achieves the intermediate representation of the embodied model, filling gaps in the intermediate representation scheme for embodied models and improving their accuracy. Moreover, since the intermediate representation information of the embodied model clearly characterizes its structure and dependencies, applying this information facilitates subsequent deployment optimization, debugging analysis, and cross-platform adaptation, enhancing the maintainability and scalability of the embodied model.
[0046] In some implementations, the sub-model includes one or more computation nodes, and the intermediate representation information of the sub-model includes a computation graph; the embodied model representation device can record the input and output information of the computation nodes used by the sub-model to construct a computation graph corresponding to the sub-model.
[0047] Please see Figure 4 In this embodiment, to more intuitively represent the sub-model, a computation graph can be used. A computation graph is a concrete representation structure of IR (Inductively Coupled Relational Model), suitable for describing models with forward propagation structures, such as CNNs, Transformers, and RNNs. Specifically, a computation graph is a directed graph used to represent a series of computational operations and data dependencies in a model. The computation graph includes nodes for representing functions (e.g., addition functions, convolution functions, activation functions) and edges for representing data flows.
[0048] In this implementation, each computation node in the sub-model represents a specific computational operation, such as tensor addition, matrix multiplication, or convolution. The relationships between computation nodes reflect the data flow structure within the sub-model.
[0049] In this embodiment, nodes in the computation graph can be generated by recording the input and output information of the computation nodes used by the sub-model. For details, please refer to [link to relevant documentation]. Figure 4 .exist Figure 4 The diagram schematically illustrates a local computation graph, which includes multiple nodes, each representing a computation function. The directed input edges of a node contain input information, and the directed output edges of a node contain output information.
[0050] In some implementations, the embodied model representation device may use a first tracking function to record the input information of the computing node before inputting the input information into the computing node's computation function; and use a second tracking function to record the output information of the computing node after the computing node's computation function has calculated the output information.
[0051] In this embodiment, to ensure that the input and output information of the computing node is traceable, a first tracking function and a second tracking function can be inserted before and after the computing function of the computing node, respectively. Optionally, the first tracking function and the second tracking function can be hook functions, decorator functions, or overloaded functions; this embodiment does not limit the specific implementation.
[0052] The first tracing function executes before the computation function is called. Specifically, it extracts the original tensor from the tensor containing tracing information output by the previous tracing function and uses it as input to the computation function. The second tracing function is called after the computation function has finished executing. Specifically, it packages the output of the computation function into a tensor containing tracing information for use by the next tracing function.
[0053] Please see Figure 5 Specifically, the first tracking function takes a tensor containing tracking information (a_traced) as input and outputs the original tensors (a, b). The tensor containing tracking information records the input information received by the computation function, including but not limited to: the original tensor, its sub-model, data type, and data source. The first tracking function extracts the original tensors (a, b) from the tensor containing tracking information (a_traced) as the actual input to the computation function, enabling the function to execute the corresponding computational logic (e.g., a+b) and output the original tensor (c). The second tracking function takes the original tensor (c) as input and a tensor containing tracking information (c_traced) as input. The second tracking function records the output information of the computation function and writes it to the computation graph. The output information includes, but is not limited to: the identifier of the output tensor and a numerical summary.
[0054] In some implementations, the embodied model representation device can apply a target encapsulation standard to the computational graph construction process of each sub-model; wherein the target encapsulation standard is used to define the recording format of the input and output information of the computational nodes.
[0055] In this embodiment, the target encapsulation standard is used to define the input and output format of each node. Optionally, the target encapsulation standard is used to define whether to record the tensor shape and data type, and whether to attach operator metadata (e.g., operation name, call location).
[0056] In this embodiment, the computational graph construction process constrained by the target encapsulation standard generates input and output information with a unified structure / form, which can be used to uniformly represent the data flow and computational logic of the sub-model.
[0057] In some implementations, the embodied model representation device can determine the intermediate representation information of each sub-model based on the segmentation markers in the embodied model and the intermediate representation information type corresponding to the model type of each sub-model; wherein the segmentation markers are used to segment multiple sub-models and / or multiple computation nodes in the embodied model.
[0058] In this embodiment, to accurately map the sub-models and intermediate representation information, segmentation markers can be added to the embodied model. These segmentation markers are structural symbols that delineate the boundaries of the sub-models; for details, please refer to [reference needed]. Figure 6For example, the segmentation marker is represented as `scope_break`. During the process of obtaining intermediate representation information of sub-models according to the order of sub-model calls in the embodied model, each computation node can be associated with the current sub-model before a segmentation marker is detected. Upon detecting a segmentation marker, each subsequent computation node can be associated with the next sub-model, until the next segmentation marker is detected.
[0059] In some implementations, segmentation tags can be added to the embodied model in various forms, such as directly callable display functions, decorators for tagging specified functions, and specified function detectors.
[0060] In some implementations, the splitting marker is added before or after a specified computation node of each sub-model in the embodied model; or, the splitting marker is added before or after a specified computation node of the splitting instruction.
[0061] In this embodiment, the segmentation marker can be inserted into the embodied model in a variety of ways, such as adding it according to the position of a specified computation node or adding it at a custom position.
[0062] In this implementation, the factor models are of different types, therefore the designated computation nodes for different sub-models can be different. A designated computation node is a computation node that delineates the boundaries of different sub-models during the execution of the embodied model. Designated computation nodes generally possess characteristics such as clear semantics, structural independence, and locationability, and can be used to mark the start or end of a sub-model.
[0063] In some implementations, for large language models built with Huggingface, the node to which `generate()` belongs can be designated as the computation node. For the ResNet visual model, the node to which `forward()` or `features()` represents the boundary of the visual feature extraction sub-model belongs can be designated as the computation node. For the YOLOv5 visual model, the node to which `detect()` / `forward_head()` converts the visual backbone output to the detection output can be designated as the computation node. For the acoustic model, the node to which `encode_audio()` represents the boundary of the speech feature extraction module belongs can be designated as the computation node. For the decoding module, the node to which `decode()` represents the entry point of the transcription sub-model belongs can be designated as the computation node. For the decision model, the node to which `select_action()` represents the boundary of the control module belongs can be designated as the computation node.
[0064] In this embodiment, the embodied model can be marked with a segmentation mark based on a specified computing node or a segmentation instruction, providing a variety of personalized segmentation methods so that the embodied model can be segmented correctly, reasonably and in accordance with user needs, thereby helping to generate more accurate intermediate representation information.
[0065] In some implementations, each sub-model includes a visual model and a large language model; the intermediate representation information of the visual model includes a computation graph in ONNX format or tvm.relax format; the intermediate representation information of the large language model includes a JSON file and model parameters; the intermediate representation information between different sub-models includes a computation graph in ONNX format, tvm.relax format, or prototext format.
[0066] In this implementation, the visual model (e.g., ResNet, YOLO, UNet) has an ordered forward propagation process: convolution, normalization, activation, pooling, and fully connected layers. Based on the clear structure and explicit tensor dependencies of the visual model, a computational graph can be used for its intermediate representation. The ONNX or tvm.relax format, as the intermediate representation information type for the visual model, can clearly and explicitly characterize the visual model. Using the ONNX or tvm.relax format to represent the visual model's computational graph offers significant advantages in model deployment, optimization, and hardware adaptation.
[0067] In this embodiment, the large language model (e.g., GPT) has a stacked structure of repetitive modules (e.g., multi-layer Transformer), where the internal state processing of each layer depends on the context, and its reasoning process includes dynamic behavior (e.g., the generation result of each step depends on the output result of the previous step), making it difficult to abstract into a computational graph. Therefore, this application uses a JSON file and model parameters for intermediate representation. Specifically, it can store the configuration file (e.g., json / yaml) of the original framework (e.g., PyTorch) and the model weight file, which allows for a more accurate intermediate representation of the large language model to accurately reproduce its generation logic.
[0068] In this embodiment, the intermediate representation information between different sub-models represents the relationship between the intermediate representation information of the sub-models. Therefore, it has a clear and orderly structure, and the corresponding computation graph can be output using any format such as ONNX, tvm.relax, or prototext. Optionally, the computation graph between different sub-models can show the details of the computation graph of each sub-model and the calling relationship between the sub-models, or it can only show the calling relationship between the sub-models. This embodiment does not limit this.
[0069] Please see Figure 7 This application also provides an embodied model representation device. The embodied model includes multiple sub-models corresponding to at least two model types; the device includes: a first intermediate representation module, used to determine the intermediate representation information of each sub-model based on the intermediate representation information type corresponding to the model type of each sub-model; the intermediate representation information types corresponding to different model types are different; a second intermediate representation module, used to determine the intermediate representation information between different sub-models based on the calculation relationship between different sub-models.
[0070] In this embodiment, the specific functions and effects of the embodied model representation device can be explained by referring to other embodiments of this application, and will not be repeated here.
[0071] Please see Figure 8 This application also provides a computer device comprising: a memory and a processor, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the method described above.
[0072] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method as described above.
[0073] This application also provides a computer program product containing instructions that, when executed by a processor, implement the method as described above.
[0074] It is understood that the specific examples in this document are only intended to help those skilled in the art better understand the embodiments of this application, and are not intended to limit the scope of the invention.
[0075] It is understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0076] It is understood that the various implementation methods described in this application can be implemented individually or in combination, and the implementation methods in this application are not limited in this respect.
[0077] Unless otherwise stated, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The term "and / or" as used in this application includes any and all combinations of one or more of the associated listed items. The singular forms "a," "the," and "the" as used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0078] It is understood that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiments can be completed by the integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0079] It is understood that the memory in the embodiments of this application may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Specifically, non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM). It should be noted that the memory in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0080] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the information to be retrieved. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0081] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the aforementioned method implementations, and will not be repeated here.
[0082] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0083] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0084] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0085] If the aforementioned function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the information to be retrieved in this application, essentially or in terms of its contribution to the prior art, or a portion of the information to be retrieved, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0086] The above description is merely a specific embodiment of this application, but the scope of protection of this invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this invention should be determined by the scope of the claims.
Claims
1. A method for representing embodied models, characterized in that, The embodied model operates within an agent equipped with sensors and actuators, and is used to realize a closed-loop motion mechanism for the agent consisting of perception, decision-making, and action. The embodied model is applied to at least one scenario in intelligent robot control, autonomous driving, and simulation training. The embodied model includes multiple sub-models corresponding to at least two model types. The multiple sub-models include a visual model and a control model. The visual model processes data collected by the visual module to perform at least one task among image recognition, object detection, and semantic segmentation. The control model generates control policies to drive the agent to perform target behaviors. The method includes: The intermediate representation information of each sub-model is determined based on the intermediate representation information type corresponding to the model type of each sub-model; the intermediate representation information types are different for different model types. Based on the computational relationships between different sub-models, intermediate representation information between different sub-models is determined; wherein, the intermediate representation information between different sub-models is used to characterize the data dependencies, input-output paths, and upstream and downstream call relationships between sub-models; The step of determining the intermediate representation information of each sub-model based on the intermediate representation information type corresponding to the model type of each sub-model includes: Based on the segmentation markers in the embodied model and the intermediate representation information type corresponding to the model type of each sub-model, the intermediate representation information of each sub-model is determined; wherein, the segmentation markers are used to segment multiple sub-models and / or multiple computation nodes in the embodied model; The sub-model includes one or more computation nodes, and the intermediate representation information of the sub-model includes a computation graph, which is a directed graph used to represent computational operations and data dependencies in the model. The computation nodes are used to represent a single computational operation. The step of determining the intermediate representation information of the sub-model includes: The input and output information of the computation nodes used by the sub-model are recorded to construct a computation graph corresponding to the sub-model; The step of recording the input and output information of the computation nodes used by the sub-model includes: Before inputting the input information into the computation function of the computing node, the input information of the computing node is recorded using the first tracking function; After the computation function of the computation node calculates the output information, the second tracking function is used to record the output information of the computation node.
2. The method according to claim 1, characterized in that, The method further includes: The target encapsulation standard is applied to the computation graph construction process of each sub-model; wherein, the target encapsulation standard is used to define the recording format of the input and output information of the computation nodes.
3. The method according to claim 1, characterized in that, The splitting marker is added before or after the specified computation node of each sub-model in the embodied model; or, the splitting marker is added before or after the specified computation node of the splitting instruction.
4. The method according to any one of claims 1 to 3, characterized in that, Each sub-model includes a visual model and a large language model; the intermediate representation information of the visual model includes a computation graph in ONNX format or tvm.relax format; the intermediate representation information of the large language model includes a JSON file and model parameters; the intermediate representation information between different sub-models is a computation graph in ONNX format, tvm.relax format, or prototext format.
5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the embodied model representation method according to any one of claims 1 to 4.
6. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the embodied model representation method according to any one of claims 1 to 4.
7. A computer program product, characterized in that, When the computer program product is executed by a processor, it implements the embodied model representation method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Data processing method and device
CN111552477A
End-to-end learning method, system and equipment based on multi-modal large model
CN118211643A