Address mapping method and device

By dynamically mapping shared physical memory to the virtual addresses of multiple chips, the problem of limited memory sharing caused by bit width differences between multiple chips is solved, and the overall performance and memory utilization of electronic devices are improved.

CN120705076AActive Publication Date: 2025-09-26HUAWEI TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202411319910.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2025-09-26
Estimated Expiration
2044-09-20

AI Technical Summary

Technical Problem

The overall performance of electronic devices has not improved much, mainly because memory sharing is limited due to differences in bit width between multiple chips, which affects the available memory size and the number of applications.

Method used

By dynamically mapping shared physical memory to the virtual addresses of multiple chips, a phased address mapping method is adopted, and flags are set to manage the mapping process of virtual addresses, thereby improving mapping efficiency and accuracy.

Benefits of technology

It improves the overall performance of electronic devices, increases the number of applications that can be opened simultaneously and memory utilization, and reduces resource usage and low interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705076A_ABST
    Figure CN120705076A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an address mapping method and device, and the method can divide a memory allocation operation into a production stage and a consumption stage, and each stage dynamically maps a virtual address for a related chip according to needs, so that the chips do not influence each other, the shared physical memory among multiple chips can be improved, and the memory allocation efficiency is improved. And the overall performance of the electronic equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of electronic devices, and in particular to an address mapping method and device. Background Art

[0002] With the rapid development of electronic devices, the performance of some hardware of electronic devices is becoming stronger and stronger. For example, the bit width of hardware such as the graphics processing unit (GPU) and central processing unit (CPU) of electronic devices is also getting larger and larger.

[0003] However, although the hardware performance of electronic devices is getting stronger and stronger, the overall performance of electronic devices has not improved much. Summary of the Invention

[0004] The present application provides an address mapping method and device, which are beneficial to improving the shared physical memory between multiple chips, thereby improving the overall performance of electronic devices.

[0005] In a first aspect, an embodiment of the present application proposes an address mapping method, which is applied to an electronic device, the electronic device includes a first application, a first device, a second device, a third device and a physical memory, the bit width of the first device, the bit width of the second device and the bit width of the third device are not completely equal, and the method may include.

[0006] The first application requests memory. The electronic device then requests a first virtual address for the first device and a second virtual address for the second device. It allocates first memory in the physical memory and maps the first virtual address and the second virtual address to a first physical address of the first memory, respectively. The first memory is used to store at least the first data written by the first device. The electronic device then requests a third virtual address for a third device and maps the third virtual address to the first physical address. The third device then retrieves the first data from the first memory corresponding to the first physical address for processing.

[0007] In an embodiment of the present application, by applying for the first virtual address of the first device in the data production stage, allocating the first memory in the physical memory, and mapping the first virtual address to the first physical address of the first memory, the first device can write the first data to the first memory, and then apply for the second virtual address of the second device, and map the second virtual address to the first physical address of the first memory respectively. In this way, if the second device acts as a producer device, it can write the second data to the first memory, and if the second device acts as a consumer device, it can also obtain the first data from the first memory for processing. Then, when the third device acts as a consumer device, it can apply for the third virtual address of the third device and map the third virtual address to the first physical address. In this way, the third device can obtain the first data from the first memory corresponding to the first physical address for processing. Since the mapping of the virtual address is divided into multiple stages, such as the data production stage and the data consumption stage, the shared physical memory can be dynamically mapped to the virtual addresses of multiple chips, thereby improving the overall performance of the electronic device.

[0008] In a possible implementation, after the first application applies for memory, the method may further include:

[0009] A first identifier of the first device and the second device is set, where the first identifier indicates that the device has been mapped to the first physical address.

[0010] In an embodiment of the present application, setting the first identifier of the first device indicates that the first device has been mapped to the first physical address. In other words, the virtual address of the first device applied for has been mapped to the first physical address. Setting the second identifier of the second device indicates that the second device has been mapped to the first physical address. In other words, the virtual address of the second device applied for has been mapped to the first physical address. Optionally, the identifier of the embodiment of the present application can also be called a flag bit.

[0011] In the embodiment of the present application, after the virtual address of the device is mapped to the physical address, an identifier is set. In this way, it can be determined through the identifier whether the physical address is mapped to the virtual address of the device, so that data can be read and written to the memory corresponding to the physical address in a timely manner, thereby improving the efficiency of data processing. In addition, the number of virtual addresses of the device applied for can be reduced.

[0012] In one possible implementation, applying for a third virtual address of a third device and mapping the third virtual address to the first physical address includes:

[0013] When it is determined that data is processed through a third device and the first identifier of the third device is not set, a third virtual address of the third device is applied for, and the third virtual address is mapped to the first physical address, and the first identifier of the third device is set, and the first identifier indicates that the device has been mapped to the first physical address.

[0014] In an embodiment of the present application, when it is determined that data is processed through a third device and the first identifier of the third device is not set, it means that the virtual address of the third device has not been mapped to the first physical address. Therefore, the third virtual address of the third device can be applied for and the third virtual address can be mapped to the first physical address. That is to say, when the third device acts as a consumer device, it is determined by the identifier whether it is necessary to apply for the virtual address of the third device and map it to the first physical address. This can improve the accuracy and efficiency of applying for the virtual address of the third device and mapping it.

[0015] In a possible implementation, before applying for a third virtual address of the third device and mapping the third virtual address to the first physical address, the method further includes:

[0016] Obtain a first set and a second set, where the first set includes devices with a first identifier set, and the second set includes devices for processing data, where the devices with the first identifier set include the first device and the second device, and the devices for processing data include the third device;

[0017] If it is determined that a first difference set exists based on the first set and the second set, applying for a virtual device of a device in the first difference set, and mapping the virtual device of the device in the first difference set to a first physical address, the device in the first difference set is a device in the second set and the device in the first difference set is different from any device in the first set;

[0018] Applying for a third virtual address of a third device and mapping the third virtual address to the first physical address includes:

[0019] In a case where the first difference set includes the third device, a third virtual address of the third device is requested, and the third virtual address is mapped to the first physical address.

[0020] In this embodiment of the present application, the devices in the first difference set are identified consumer devices and are not marked with the first identifier. In other words, the devices in the first difference set need to apply for a virtual address and map it to the first physical address. If the first difference set includes a third device, it means that the third device needs to apply for a virtual address and map it to the first physical address. In this case, a third virtual address of the third device is applied for and mapped to the first physical address.

[0021] In a possible implementation, the electronic device further includes a first service, and the first application applies for memory, including:

[0022] The first application requests memory from the first service;

[0023] Applying for a first virtual address of a first device and applying for a second virtual address of a second device, allocating a first memory in a physical memory, and mapping the first virtual address and the second virtual address to a first physical address of the first memory, respectively, includes:

[0024] The first service applies for a first virtual address of the first device and a second virtual address of the second device, allocates a first memory in the physical memory, and maps the first virtual address and the second virtual address to a first physical address of the first memory respectively.

[0025] In an embodiment of the present application, the first application can apply for memory from the first service, and the first service can apply for the first virtual address of the first device and the second virtual address of the second device, allocate the first memory in the physical memory, and map the first virtual address and the second virtual address to the first physical address of the first memory respectively. In this way, the problems of inefficiency and high resource usage caused by the interaction of multiple services to achieve address mapping can be reduced, thereby improving the efficiency of address mapping and reducing resource usage.

[0026] In one possible implementation, applying for a third virtual address of a third device and mapping the third virtual address to the first physical address includes:

[0027] The first service applies for a third virtual address of the third device and maps the third virtual address to the first physical address.

[0028] In an embodiment of the present application, the first service can also apply for the virtual address of the third device and perform mapping. This can reduce the inefficiency and high resource usage caused by the interaction of multiple services to achieve address mapping, thereby improving the efficiency of address mapping and reducing resource usage.

[0029] In a possible implementation, the first application applies for memory, including:

[0030] In response to opening a window on the electronic device, the first application requests a graphics memory, and the graphics memory is used to store graphics data required by the window.

[0031] In an embodiment of the present application, the first application requests graphics memory in response to opening a window on the electronic device, which can increase the number of windows opened on the electronic device.

[0032] In one possible implementation, the first memory is also used to store second data written by the second device, that is, the second device writes data to the first memory as a producer device.

[0033] In another possible implementation, applying for the second virtual address of the second device includes:

[0034] When it is determined that data is processed by the second device and the first identifier of the second device is not set, a second virtual address of the second device is applied for and the first identifier of the second device is set, where the first identifier indicates that the device has been mapped to the first physical address.

[0035] In the embodiment of the present application, the second device serves as a consumer device. If it is determined that the second device is required and the first identifier of the second device is not set, it means that the virtual address of the second device is not mapped to the first physical address, and the second device has difficulty accessing the first memory normally. Therefore, the second virtual address of the second device is requested, so that the second virtual address can be mapped to the first physical address, allowing the second device to access the first memory normally, and then obtain the first data in the first memory for processing.

[0036] In a second aspect, embodiments of the present application provide an address mapping device, comprising a processor coupled to a memory and configured to execute instructions in the memory to implement the method of any possible implementation of the first aspect. Optionally, the device further comprises a memory. Optionally, the device further comprises a communication interface, the processor coupled to the communication interface.

[0037] In a third aspect, an embodiment of the present application provides a processor, comprising: an input circuit, an output circuit, and a processing circuit. The processing circuit is configured to receive a signal through the input circuit and transmit a signal through the output circuit, so that the processor executes the method of any possible implementation of the first aspect.

[0038] In a specific implementation, the processor may be a chip, the input circuit may be an input pin, the output circuit may be an output pin, and the processing circuit may be a transistor, a gate circuit, a trigger, or various logic circuits. The input signal received by the input circuit may be, for example, but not limited to, received and input by a receiver, and the signal output by the output circuit may be, for example, but not limited to, output to and transmitted by a transmitter. The input circuit and the output circuit may be the same circuit, which functions as an input circuit and an output circuit at different times. The embodiments of the present application do not limit the specific implementation of the processor and various circuits.

[0039] In a fourth aspect, an address mapping device is provided, comprising a processor and a memory. The processor is configured to read instructions stored in the memory and receive signals via a receiver and transmit signals via a transmitter to execute the method of any possible implementation of the first aspect.

[0040] Optionally, there are one or more processors and one or more memories.

[0041] Optionally, the memory may be integrated with the processor, or the memory may be provided separately from the processor.

[0042] In the specific implementation process, the memory can be a non-transitory memory, such as a read-only memory (ROM), which can be integrated with the processor on the same chip or can be set on different chips. The embodiments of the present application do not limit the type of memory and the setting method of the memory and the processor.

[0043] It should be understood that related data interaction processes, such as sending indication information, can be the process of outputting indication information from the processor, and receiving capability information can be the process of receiving input capability information from the processor. Specifically, the output data of the processing can be output to the transmitter, and the input data received by the processor can come from the receiver. The transmitter and receiver can be collectively referred to as a transceiver.

[0044] The address mapping device in the fourth aspect mentioned above can be a chip, and the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading the software code stored in the memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.

[0045] In a fifth aspect, a computer program product is provided, which includes: a computer program (also referred to as code, or instructions), which, when executed, enables a computer to execute the method in any possible implementation of the first aspect.

[0046] In a sixth aspect, a computer-readable storage medium is provided, which stores a computer program (also referred to as code, or instructions) which, when run on a computer, enables the computer to execute the method in any possible implementation of the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 A schematic diagram of a scenario in which multiple chips perform data processing according to an embodiment of the present application;

[0048] Figure 2 A schematic diagram illustrating the limitation of the memory size that can be used by an electronic device as a whole in the related art;

[0049] Figure 3 A schematic diagram of the hardware architecture of an electronic device provided in an embodiment of the present application;

[0050] Figure 4 A schematic diagram of the software architecture of an electronic device provided in an embodiment of the present application;

[0051] Figure 5 A schematic diagram of a data processing system architecture provided in an embodiment of the present application;

[0052] Figure 6 A flowchart of a method for dynamically mapping shared physical memory to virtual addresses of multiple devices provided in an embodiment of the present application;

[0053] Figure 7 A flowchart of another method for dynamically mapping shared physical memory to virtual addresses of multiple devices proposed in an embodiment of the present application;

[0054] Figure 8 A schematic diagram of physical memory usage provided in an embodiment of the present application;

[0055] Figure 9 A flowchart of an address mapping method provided in an embodiment of the present application;

[0056] Figure 10 A schematic diagram of the framework of an address mapping device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0057] In order to make the purpose and technical solution of this application clearer and more intuitive, the following will be combined with the accompanying drawings and embodiments to describe the embodiments of this application in detail through the address mapping method and device. It should be understood that the specific embodiments described here are only used to explain this application and are not used to limit this application.

[0058] Before introducing the methods and devices provided in the embodiments of the present application, the following points are explained.

[0059] First, in the embodiments described below, the various terms and abbreviations, such as "CPU," are provided for ease of description and should not constitute any limitation on this application. This application does not exclude the possibility of defining other terms in existing or future protocols that can achieve the same or similar functions.

[0060] Second, the first, second and various numerical numbers in the embodiments shown below are only used for the convenience of description and are not intended to limit the scope of the embodiments of the present application.

[0061] Third, "at least one" means one or more, and "more" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b and c can mean: a, or b, or c, or a and b, or a and c, or b and c, or a, b and c, where a, b, c can be single or multiple.

[0062] The technical solution in this application will be described below with reference to the accompanying drawings.

[0063] The technical solutions of the embodiments of the present application can be applied to various communication systems, such as: long term evolution (LTE) system, LTE frequency division duplex (FDD) system, LTE time division duplex (TDD), universal mobile telecommunication system (UMTS), fifth generation (5G) system or new radio (NR) or other evolved communication systems.

[0064] The terminal device in the embodiments of the present application may also be referred to as: user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, device, wireless communication device, electronic device, user agent or user device, etc.

[0065] The terminal device may be a device that provides voice / data connectivity to users, such as a handheld device or vehicle-mounted device with wireless connection function. At present, some examples of terminals include: mobile phones, car computers, tablet computers, laptop computers, PDAs, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to wireless modems, vehicle-mounted devices, wearable devices, terminal devices in 5G networks or future evolved public land mobile communication networks (PLMNs). The terminal equipment in the network (PLMN), etc., is not limited to this in the embodiments of the present application.

[0066] As an example and not a limitation, in the embodiment of the present application, the terminal device may also be a wearable device. Wearable devices may also be called wearable smart devices, which are a general term for wearable devices that are intelligently designed and developed using wearable technology for daily wear, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothes or accessories. Wearable devices are not only hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are fully functional, large in size, and can achieve complete or partial functions without relying on smartphones, such as smart watches or smart glasses, as well as those that only focus on a certain type of application function and need to be used in conjunction with other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0067] In addition, in the embodiments of the present application, the terminal device can also be a terminal device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection. The terminal device of the present application can also be an on-board unit, on-board module, on-board component, on-board chip or on-board unit built into the vehicle as one or more components or units. The vehicle can implement the method of the present application through the built-in on-board unit, on-board module, on-board component, on-board chip or on-board unit. Therefore, the embodiments of the present application can be applied to vehicle networks, such as vehicle to everything (V2X), long-term evolution-vehicle (LTE-V), vehicle-to-vehicle (V2V), etc.

[0068] At present, although the performance of some hardware of electronic devices is getting stronger and stronger, for example, the bit width of hardware such as GPU and central processing unit of electronic devices is getting larger and larger, the overall performance of electronic devices has not been greatly improved. In one possible scenario, the reasons for the small performance improvement of electronic devices as a whole include: the hardware with the lowest bit width will become a short board, resulting in a barrel effect, affecting the overall available memory size, resulting in a small improvement in the overall performance of electronic devices. For example, the number of web pages that a browser can open is limited. Once the number of web pages opened in the browser is too large, it will cause the browser to crash or prohibit the opening of new web pages.

[0069] The following uses a scenario where a CPU, GPU, and display subsystem (DSS) chip are used for data processing to illustrate why the lowest bit width becomes a shortcoming, affecting the overall available memory size.

[0070] See also Figure 1 , Figure 1 This is a schematic diagram of a scenario in which multiple chips perform data processing according to an embodiment of the present application. Figure 1 The illustrated scenario may include a CPU, GPU, DSS, and physical memory. The CPU and GPU may have a bit width of 64 bits, and the DSS may have a bit width of 32 bits. Physical memory may include a buffer, which may reside in random access memory (RAM) or a buffer area on a hard disk.

[0071] In an embodiment of the present application, the buffer needs to be accessible to the CPU, GPU, and DSS, and the memory of the buffer is physically shared, so that data can be transmitted between the CPU, GPU, and DSS through the buffer. Taking data processing including image processing as an example, the original layer drawing can be completed by the CPU or GPU first, which can also be understood as the production of data by the CPU or GPU; and then the image to be displayed is synthesized by the GPU or DSS, which can also be understood as the consumption of data by the GPU or DSS. In an embodiment of the present application, the CPU can draw layers, and the GPU can write to the disk (flush) and synthesize (compose) images. Flush generally refers to the operation of refreshing buffered data from a cache or temporary memory to a permanent memory or an external device. Compose is generally related to the creation, combination, or synthesis of images. It can refer to the combination of multiple image elements, layers, or image fragments into a complete image.

[0072] Optionally, scenarios for performing image processing using a CPU, a GPU, and a DSS include but are not limited to the following scenarios:

[0073] In the data production phase, the CPU and / or GPU completes the original layer drawing and writes the drawn data into the buffer. Then, in the data consumption phase, the GPU and / or DSS performs subsequent data processing, such as writing the drawn data into read-only memory (ROM) or an external device, or using the drawn data to synthesize the image to be displayed.

[0074] In the embodiments of the present application, the CPU completing the drawing of the original layer can be understood as production data, while the GPU performing image enhancement processing such as rendering on the drawn data can be understood as consumption data. Furthermore, the GPU performing image enhancement processing such as rendering on the drawn data can also be understood as production data, while the DSS using this image enhancement data to synthesize the image to be displayed can be understood as consumption data.

[0075] It should be understood that Figure 1 The scenario shown can also be applied to other data processing, such as training or reasoning of AI models. The specific data processing is not limited here. Training or reasoning of AI models can be compared to the image processing processes performed by CPUs, GPUs, and DSSs. The differences between training or reasoning AI models and image processing include the use of different chips and processes.

[0076] In related technologies, when data processing is required, a virtual address is first requested in the process of the application used to draw the layer. The application can also be referred to as an application or program. When the program wants to read or write the content of this address, the system allocates a memory page from the physical memory and then maps the virtual address to the physical address.

[0077] Taking the above scenario as an example, it is necessary to first apply for virtual addresses in the CPU, GPU and DSS respectively, and then map the virtual addresses applied by the CPU, GPU and DSS to the same physical address, so as to achieve shared processing between the CPU, GPU and DSS. It can also be understood as achieving the processing of the production data stage and the consumption data stage between the CPU, GPU and DSS. The virtual address range is affected by the chip bit width n. For example, the maximum range of the virtual address is 2 n , and as for the physical address range, it is related to the capacity of physical memory.

[0078] From this, we can see that since virtual addresses are first applied for separately in the CPU, GPU, and DSS, and then the virtual addresses applied for by the CPU, GPU, and DSS are mapped to the same physical address, thereby achieving shared processing among the CPU, GPU, and DSS, the available physical address range is affected by the chip with the lowest bit width among the three chips: CPU, GPU, and DSS. In other words, the chip with the lowest bit width among the three chips: CPU, GPU, and DSS will become the weak link, resulting in a "barrel effect" and affecting the overall available memory size. Taking the number of web pages opened by a browser as an example, since image elements need to be displayed when opening a web page, physical memory needs to be applied for to store the image elements. However, when the overall available memory size of the physical address is limited, it will affect the number of web pages opened by the browser.

[0079] See also Figure 2 , Figure 2 This is a schematic diagram showing the limitation of the memory size that can be used by the entire electronic device in the related art. Figure 2 As shown, since the bit width of CPU and GPU are both 64 bits, the maximum memory size that CPU and GPU can use theoretically is 2 64 =8 gigabytes (G), while the maximum memory size that DSS can theoretically use is 2 32=4G. Even if an electronic device has 16G or even 32G of physical RAM, and the CPU and GPU have a 64-bit bit width, due to the bit width limitation of the DSS, the CPU, GPU, and DSS can only use 4G of memory. The excess memory cannot be used for graphics memory. For example, for 16G of RAM, only 4G of memory can be shared by the CPU, GPU, and DSS as graphics memory, leaving 12G of remaining memory unavailable for graphics. This will result in significant limitations on the number of applications that can be opened simultaneously on the electronic device and the performance of the electronic device. This shows that even if the performance of some hardware in an electronic device is getting stronger, the overall performance of the electronic device will not improve much.

[0080] In general, when multiple chips share physical memory, the total available memory is limited by the chip with the smallest bit width. For example, if a 32-bit chip is limited to 4GB of physical memory, any memory above 4GB cannot be used, resulting in wasted memory. For 16GB of physical memory, the utilization rate is only 4GB / 16GB = 25%, and for 32GB of physical memory, the utilization rate is only 4GB / 32GB = 12.5%.

[0081] In view of this, embodiments of the present application provide an address mapping method and apparatus that can improve the overall performance of electronic devices by dynamically mapping shared physical memory to the virtual addresses of multiple chips. Taking the number of web pages that can be opened by a browser as an example, the shared memory that consumes data with chips that require a minimum bit width is limited by the minimum bit width, but the shared memory that consumes data with chips that do not require a minimum bit width is not limited by the minimum bit width. This doubles the number of new applications that can be opened by the entire machine, allowing more new web pages or applications to be opened as long as there is physical memory.

[0082] The address mapping method is described below in conjunction with the hardware architecture and software architecture of the electronic device.

[0083] See also Figure 3 , Figure 3 A schematic diagram of the hardware architecture of an electronic device provided in an embodiment of the present application.

[0084] like Figure 3As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0085] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0086] The processor 110 may include one or more processing units, for example, an application processor (AP), a modem processor, a CPU, a GPU, a DSS, an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0087] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.

[0088] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly retrieve it from the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0089] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0090] The internal memory 121 can be used to store computer executable program codes, and the executable program codes include instructions. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 100 (such as audio data, image data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running instructions stored in the internal memory 121 and / or instructions stored in a memory provided in the processor.

[0091] In an embodiment of the present application, taking data processing including image data (also referred to as graphic data) processing as an example, the drawing of the layer can be completed by the CPU or GPU, and the drawn image data can be stored in the storage data area. The drawn image data is then obtained from the storage data area by the GPU or DSS for image synthesis, and the synthesized image data is then displayed through the display screen 194. Taking the training or reasoning of an artificial intelligence model as an example, the user's operation data can be obtained by the CPU, and the operation data can be stored in the storage data area, and then the operation data can be obtained from the storage data area by the NPU for training or reasoning. Among them, the operation data can be data generated by detecting the user's operation on the electronic device, such as data of an application clicked by the user detected by the electronic device.

[0092] See also Figure 4 , Figure 4A schematic diagram of the software architecture of an electronic device provided in an embodiment of the present application.

[0093] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present invention, the software structure of the electronic device 100 is exemplified by the Harmony system of the layered architecture.

[0094] The Harmony system adopts a multi-kernel design, optionally including the Linux kernel, the Harmony microkernel and LiteOS. Through this design, devices with different device capabilities can choose the appropriate system kernel. The kernel layer also includes a kernel abstraction layer (Kernal Abstract Layer), which provides basic kernel capabilities to other Harmony layers, such as process management, thread management, memory management, file system management, network management and peripheral management. Optionally, the kernel layer also includes a rendering service (render service), which is generally responsible for converting user interface (UI) elements or graphic data into images that can be displayed on the screen. This includes processing visual effects such as color, texture, lighting, shadows, and making graphic elements be drawn to the screen in the correct manner and order to achieve smooth animation and interactive experience. It should be understood that the rendering service can also be configured at other levels of the software framework, and is not limited to the kernel layer, as long as the rendering service can achieve its functions in the embodiments of the present application.

[0095] The system basic service layer is the core capability set of the Harmony system, which supports the Harmony system to provide services to application services through the framework layer in the scenario of multi-device deployment. This layer optionally includes the following parts:

[0096] The system's basic capability subsystems provide fundamental capabilities for running, scheduling, and migrating distributed applications across multiple devices in the Harmony system. These subsystems comprise a distributed soft bus, distributed data and file management, distributed task scheduling, the Ark runtime, and distributed security and privacy protection. The Ark runtime provides a C / C++ / JavaScript multi-language runtime and basic system class libraries. It also provides a runtime for Java programs statically compiled using the Ark compiler (i.e., applications or frameworks developed in Java).

[0097] Basic software service subsystems: These provide common, general-purpose software services for the Harmony system and consist of subsystems such as graphics and imaging, distributed media, distributed AI, multi-mode input, MSDP & DV, event notification, phone services, and distributed DFX. These subsystems can be tailored to the specific functional granularity within each subsystem, tailored to the deployment environment of different device form factors.

[0098] Enhanced Software Service Subsystem Set: This provides the Harmony system with differentiated, device-specific, capability-enhancing software services. This set comprises subsystems such as tablet software, smart screen software, vehicle-mounted software, and IoT software. This enhanced software service subsystem set can be tailored to the deployment environment of different device form factors at the subsystem level, and each subsystem can be tailored to the functional level within it.

[0099] Harmony Driver Framework (HDF) and Hardware Abstraction Adaptation Layer (HAL): They are the foundation of the open hardware ecosystem of the Harmony system, providing hardware capability abstraction to the hardware upward and providing a development framework and operating environment for various peripheral drivers downward.

[0100] Hardware Service Subsystem Set: This provides common, adaptive hardware services for the Harmony system and consists of hardware service subsystems such as general sensors, location, power, USB, and biometrics. The hardware service subsystem set can be tailored to the deployment environment of different device form factors, and each subsystem can be tailored to the functional granularity.

[0101] Proprietary Hardware Service Subsystem: This subsystem provides differentiated hardware services for different devices within the Harmony system. This includes options for tablet, vehicle, wearable, and IoT hardware services. The Proprietary Hardware Service Subsystem can be tailored to the subsystem level, and each subsystem can be tailored to the functional level.

[0102] The framework layer provides Harmony system applications with a user program framework and meta-capability framework in multiple languages, including Java, C, C++, and JavaScript. It also provides a multi-language framework API for various software and hardware services. The application layer includes system applications and third-party applications, including browsers, cameras, galleries, calendars, calls, maps, navigation, Wi-Fi, Bluetooth, music, video, and short messaging. Applications in the Harmony system are built on AA and FA.

[0103] In an embodiment of the present application, physical memory is first allocated to a producer device (e.g., a CPU or GPU), and the producer device can write the produced data (e.g., drawn image data) into the physical memory. Then, when the produced data needs to be consumed, the consumer device (e.g., a GPU or DSS) used to consume the data can use the mapping between the virtual address of the consumer device applied for and the physical address where the produced data is written to obtain the produced data from the physical address where the produced data is written, and consume the produced data. The consumed data can, for example, be an image to be displayed synthesized using image data.

[0104] The producer device may be a device that produces data, and the consumer device may be a device that consumes data. The definitions of production data and consumption data can be found in the description of the above embodiments and are not elaborated here. For example, the producer device may be a CPU or a GPU, and the consumer device may be a GPU or a DSS.

[0105] To better understand the data processing scheme of the present invention, the following examples illustrate the physical address occupation during the data production and data consumption phases, using producer devices, consumer devices, and a graphics memory management decision device as examples of several functional units in an electronic device. The graphics memory management decision device can be used to manage graphics memory, such as managing the independent graphics memory of each chip.

[0106] See also Figure 5 , Figure 5 This is a schematic diagram of a data processing system architecture provided in an embodiment of the present application. Figure 5 The system architecture shown may include a producer device, a consumer device, a graphics memory management decision device and a physical memory device. Among them, the physical memory device is also referred to as physical memory. Optionally, the producer device may be one or more independent devices. When there are multiple producer devices, there may be 1-M independent producer devices, where M is an integer not less than 2. Optionally, the consumer device may be one or more independent devices. When there are multiple consumer devices, there may be 1-N independent consumer devices. In an embodiment of the present application, the producer device and the consumer device may be hardware devices, and the graphics memory management decision device may be a software device. Taking graphics processing as an example, the producer device may be used to draw and write graphics data, and the consumer device may be used to synthesize images or read graphics data for display.

[0107] In the data production stage, the producer device applies for the virtual address of the producer device. Then, the graphics memory management decision device allocates shared physical memory in the physical memory device to the producer device and marks the producer device with an identifier (also known as a flag bit). The identifier is used to indicate the physical address of the shared physical memory allocated to the producer device. In other words, the identifier is used to indicate that the virtual address of the producer device has been mapped to the physical address. In addition, the graphics memory management decision device maps the virtual address of the producer device to the physical address. The producer device can then use the mapping between the virtual address of the producer device and the physical address to write the data produced by the producer device into the memory corresponding to the physical address.

[0108] Then, in the data consumption phase, in the embodiment of the present application, in the data consumption phase, the graphics memory management decision device may perform but is not limited to the following steps:

[0109] 1. Find all consumer devices of a certain memory according to business logic.

[0110] 2. Find the difference between the device identifier set marked on the memory chip and the consumer device, and apply for a virtual address from the consumer in the difference set.

[0111] 3. Bind the newly applied virtual address to the original physical memory.

[0112] Among them, the difference set of the embodiment of the present application includes consumer devices that are not marked. For example, the graphics memory management decision device decides a set of consumer devices for consuming production data based on business logic, and the consumer device set includes one or more consumer devices. For example, the consumer device set includes n consumer devices, 0<n<N, and n is an integer. Then, the consumer device set is compared with the set of marked producer devices to determine whether there is a difference set. If there is a difference set, it means that at least one consumer device in the consumer device set is not marked, that is, at least one consumer device in the consumer device set has not applied for mapping to the virtual address of the physical address where the production data is stored. In other words, at most some of the consumer devices in the consumer device set have been marked before the data consumption stage, that is, at most some of the consumer devices in the consumer device set have applied for mapping to the virtual address of the physical address where the production data is stored.

[0113] Therefore, the graphics memory management decision device can request a virtual address from each of the at least one consumer device. The graphics memory management decision device then maps the virtual address requested by each of the at least one consumer device to the physical address. Each of the at least one consumer device can then, based on the mapping between the requested virtual address and the physical address, retrieve production data from the memory corresponding to the physical address for consumption. For identified consumer devices, these consumer devices can utilize the mapping between the requested virtual address and the physical address to retrieve production data from the memory corresponding to the physical address for consumption.

[0114] It should be noted that if there is no difference set, it means that each consumer device in the consumer device set has applied for a virtual address mapped to the physical address before consuming the data. Then each consumer device in the consumer device set can use the mapping between the applied virtual address and the physical address to obtain the production data from the memory corresponding to the physical address for consumption.

[0115] In general, the embodiments of this application divide memory allocation operations into a production phase and a consumption phase. In each phase, virtual addresses are dynamically mapped only to the relevant chips on demand, ensuring that the chips do not interfere with each other. This way, the memory available to the entire machine is no longer limited by the bit width of a single chip. The theoretical maximum is the sum of the virtual space of all chips, which is greater than the capacity of the physical address space, thus maximizing the utilization of physical memory.

[0116] For example, let's take data processing, including image data processing, as an example. Assume that the system framework includes a CPU, a GPU, and a DSS. The GPU can be both a producer and a consumer, with the CPU acting as the producer and the DSS acting as the consumer. During the data production phase, the CPU and GPU each request a virtual address. Assume the CPU requests virtual address 1 and the GPU requests virtual address 2. The graphics memory management decision module then allocates memory 1 to the CPU and GPU, maps virtual address 1 and virtual address 2 to physical address 1 of memory 1, and tags the CPU and GPU with identifiers indicating physical address 1. After producing data, the CPU can use the mapping between virtual address 1 and physical address 1 to write the data produced by the CPU into memory 1 corresponding to physical address 1. Then, during the data consumption phase, the graphics memory management decision module determines, based on business logic, a set of consumer devices to consume the produced data. This set of consumer devices includes the GPU and the DSS. Consequently, by comparing the set of consumer devices with the set of tagged producer devices, a difference set is obtained, which includes the DSS. Then, for the DSS, the graphics memory management decision device can request virtual address 3 from the DSS, and then the graphics memory management decision device maps virtual address 3 to physical address 1. The DSS can then use the mapping between virtual address 3 and physical address 1 to obtain the data generated by the CPU from memory 1 corresponding to physical address 1 and consume the generated data. For the GPU, since the GPU has been marked during the data generation phase, that is, the virtual address 2 requested by the GPU has been mapped to physical address 1 during the data generation phase, the GPU can use the mapping between virtual address 2 and physical address 1 to obtain the generated data from memory 1 corresponding to physical address 1 for consumption during the data consumption phase.

[0117] It should be noted that if the determined consumer device set includes a GPU but not a DSS, then there will be no difference when comparing the consumer device set with the set of marked producer devices. If the determined consumer device set includes a DSS but not a GPU, then there will be a difference when comparing the consumer device set with the set of marked producer devices.

[0118] In an embodiment of the present application, in the data production stage, the CPU and GPU each apply for a virtual address, and map the virtual addresses of the CPU and GPU to the physical addresses of the allocated memory. In the data consumption stage, if the GPU needs to consume data, since the GPU has already applied for a virtual address and mapped it to a physical address, the GPU no longer needs to apply for and map the virtual address in the data consumption stage, thereby improving the efficiency of data consumption. In other words, in the data production stage, the embodiment of the present application applies for the virtual addresses of multiple devices and maps the virtual addresses of multiple devices to physical addresses. If a device that has applied for a virtual address and mapped the physical address needs to consume data, then in the data consumption stage, the device that has applied for a virtual address and mapped the physical address can directly use the mapping between the virtual address and the physical address to obtain data for consumption, thereby improving the efficiency of data consumption.

[0119] It should be noted that in an embodiment of the present application, the devices involved in data processing include at least three devices, and the bit width of at least two of the at least three devices is greater than the bit width of the other devices in the at least three devices except for the at least two devices. Exemplarily, the devices involved in image processing include a CPU, a GPU, and a DSS, wherein the bit width of the CPU and the GPU is greater than the bit width of the DSS. However, the bit widths of the CPU and the GPU can be the same or different, depending on the actual situation. For example, the bit width of the CPU is greater than the bit width of the GPU, or the bit width of the CPU is equal to the bit width of the GPU, or the bit width of the CPU is less than the bit width of the GPU.

[0120] In an embodiment of the present application, in the data production stage, in addition to the producer device needing to apply for a virtual address and map it to a physical address, a virtual address can also be applied for at least one device with the largest bit width among the at least three devices, and then the virtual address of each of the at least one device with the largest bit width is mapped to the allocated physical address, and the range of physical addresses supported by the device with the largest bit width is greater than or equal to the range of physical addresses supported by devices with other bit widths, that is, the range of physical addresses supported by the device with the largest bit width can cover the range of physical addresses supported by devices with other bit widths. In this way, not only can the efficiency of data consumption be improved, but also the situation where the physical address of the larger bit width device storing production data cannot be accessed by the smaller bit width device due to the smaller bit width of the device applying for the virtual address in the data production stage can be reduced, thereby reducing the situation where data cannot be consumed due to the smaller bit width device being unable to access production data.

[0121] For example, assuming the CPU and GPU bit widths are 64 bits, and the DSS bit width is 32 bits, since the CPU and GPU bit widths are both at their maximum, then during the data production phase, the CPU and / or GPU may apply for a virtual address and map the applied virtual address to a physical address. For another example, assuming the CPU bit width is 128 bits, the GPU bit width is 64 bits, and the DSS bit width is 32 bits, since the CPU bit width is at its maximum, then during the data production phase, the CPU may apply for a virtual address and map the applied virtual address to a physical address. In this case, if the GPU is a producer device, the GPU may also apply for a virtual address and map it to a physical address. If the GPU is not a producer device, then the GPU virtual address may not be applied for first. As another example, assuming the CPU bit width is 64 bits, the GPU bit width is 128 bits, and the DSS bit width is 32 bits, since the GPU bit width is the largest, then during the data production stage, the GPU may apply for an address and map the requested virtual address to a physical address; at this time, if the CPU is a producer device, the CPU may also apply for a virtual address and map it to a physical address; if the CPU is not a producer device, the CPU virtual address may not be applied for first. As another example, assuming the CPU bit width is 32 bits, the GPU and DSS bit widths are 64 bits, since the GPU and DSS bit widths are both the largest, then during the data production stage, the GPU and / or DSS may apply for a virtual address and map the requested virtual address to a physical address. As another example, assuming the CPU bit width is 64 bits, the CPU bit width is 32 bits, and the DSS bit width is 128 bits, then during the data production stage, the DSS may apply for a virtual address and map the requested virtual address to a physical address.

[0122] It should be understood that if, during the data production phase, the producer device and the device with the maximum bit width are not the same device, the producer device needs to apply for a virtual address and map it to a physical address because it needs to store the produced data. For the device with the maximum bit width, a virtual address can also be applied for and mapped to the physical address where the produced data is stored during the data generation phase. This facilitates the device with the maximum bit width to obtain and consume the produced data in a timely manner when it is used as a consumer device. Alternatively, the device with the maximum bit width can apply for a virtual address and map it to the physical address where the produced data is stored only when it is determined to be a producer device or a consumer device.

[0123] In another possible implementation, virtual addresses can also be applied for the devices other than the device with the smallest bit width among the at least three devices, and then the virtual addresses of the devices other than the device with the smallest bit width are mapped to the allocated physical addresses. In this way, when the devices other than the device with the smallest bit width need to consume data, the data can be consumed in a timely manner, thereby improving the efficiency of consuming data.

[0124] Exemplarily, assuming that the CPU bit width is 128 bits, the GPU bit width is 64 bits, and the DSS bit width is 32 bits, then the CPU and / or GPU may apply for a virtual address and map the applied virtual address to a physical address. Furthermore, assuming that the CPU bit width is 64 bits, the GPU bit width is 128 bits, and the DSS bit width is 32 bits, then the CPU and / or GPU may apply for a virtual address and map the applied virtual address to a physical address. Furthermore, assuming that the CPU bit width is 64 bits, the GPU bit width is 32 bits, and the DSS bit width is 128 bits, then the CPU and / or DSS may apply for a virtual address and map the applied virtual address to a physical address.

[0125] In another possible implementation, it is also possible that when the device is needed to process data, it is necessary to apply for a virtual address from the device and map the virtual address to a physical address. In this way, the situation where the virtual address is applied for and mapped to the physical address but has not been used can be reduced, thereby improving the utilization rate of device resources.

[0126] It should be noted that in an embodiment of the present application, when the producer devices are marked in the data production stage, M producer devices can be marked separately, or some of the M producer devices can be marked. These some producer devices can be, for example, producer devices that actually participate in data processing in this round of data processing, such as producer devices that actually participate in drawing layer data.

[0127] Below, the method flow for dynamically mapping shared physical memory to virtual addresses of multiple devices is described. The following embodiment is described by taking multiple devices including a CPU, a GPU, and a DSS as an example. In the embodiment of the present application, the CPU and the GPU can be used as producer devices, and the GPU and the DSS can be used as consumer devices, that is, the GPU can be used as both a producer device and a consumer device. Optionally, the bit width of the CPU is greater than the bit width of the DSS, and the bit width of the GPU is greater than the bit width of the DSS, and the bit width of the CPU and the bit width of the GPU can be the same or different. For example, the bit width of the CPU and the GPU are both 64 bits, and the bit width of the DSS is 32 bits. For another example, the bit width of the CPU is 256 bits, the bit width of the GPU is 128 bits, and the bit width of the DSS is 64 bits.

[0128] See also Figure 6 , Figure 6 This is a flow chart of a method for dynamically mapping shared physical memory to multiple device virtual addresses provided in an embodiment of the present application. The method of this embodiment may include:

[0129] S611: The application requests graphics memory from the rendering service.

[0130] The graphics memory may be a space for storing graphics data, and the graphics memory corresponds to a physical address. In the embodiment of the present application, the application may be an application capable of generating graphics data, such as a camera application, a browser application, etc., which is not limited here.

[0131] S612: The rendering service applies for virtual addresses from the CPU and GPU respectively, allocates memory 1, maps the virtual address 1 applied by the CPU and the virtual address 2 applied by the GPU to the physical address 1 of the memory 1 respectively, and sets the flag bits of the CPU and GPU.

[0132] In an embodiment of the present application, the rendering service acts as a graphics memory management decision-making device. It should be understood that the functionality of the graphics memory management decision-making device can also be implemented through other services or modules, not just the rendering service. Optionally, the rendering service can transmit virtual address request requests to the CPU and GPU, respectively. Upon receiving the virtual address request, the CPU requests virtual address 1, and transmits the requested virtual address 1 to the rendering service. The rendering service can then map virtual address 1 to physical address 1, obtaining mapping relationship 1, which maps the relationship between virtual address 1 and physical address 1. The rendering service then transmits mapping relationship 1 to the CPU. Furthermore, upon receiving the virtual address request, the GPU requests virtual address 2, and transmits the requested virtual address 2 to the rendering service. The rendering service can then map virtual address 2 to physical address 1, obtaining mapping relationship 2, which maps the relationship between virtual address 2 and physical address 1. The rendering service then transmits mapping relationship 2 to the GPU. Furthermore, the rendering service sets flags for the CPU and GPU, which record that both the CPU and GPU have allocated physical address 1 and mapped it.

[0133] It should be noted that in the embodiments of the present application, the application and the rendering service can be software processes running on a CPU. The CPU running the application and the rendering service can be a different CPU from the CPU writing the layer data to memory 1 corresponding to physical address 1. In other words, the CPU running the application and the rendering service is CPU 1, and the CPU writing the layer data to memory 1 corresponding to physical address 1 is CPU 2.

[0134] S613. The CPU writes the layer data generated by the application into memory 1 based on the mapping between virtual address 1 and physical address 1.

[0135] In an embodiment of the present application, the layer data generated by the application can be transmitted to the CPU, and the CPU can write the layer data generated by the application into the memory 1 corresponding to the physical address 1, or the CPU can write the layer data generated by the application into the memory 1 corresponding to the physical address 1 after performing certain processing. It should be noted that the layer data can be drawn by the CPU running the application, or drawn by the CPU and GPU running the application together, and there is no limitation here. If the layer data is drawn by the CPU and GPU running the application together, the method of the embodiment of the present application can also include the GPU writing the layer data generated by the GPU into the memory 1 corresponding to the physical address 1 based on the mapping between the virtual address 2 and the physical address 1.

[0136] S614: The rendering service determines that the layer data needs to be synthesized by the GPU.

[0137] In the embodiment of the present application, since the GPU flag has been set in S612, the rendering service can know through the GPU flag that the virtual address requested by the GPU has been mapped to the physical address 1. In other words, the physical address 1 already has the GPU flag, so the rendering service does not need to find the GPU to apply for the virtual address.

[0138] It should be noted that if the GPU does not have a flag bit, it is necessary to apply for a virtual address from the GPU, and then map the virtual address 2 applied by the GPU to the physical address 1.

[0139] S615: The rendering service instructs the GPU to synthesize an image.

[0140] In an embodiment of the present application, the rendering service may transmit an indication message 1 to the GPU, where the indication message 1 is used to instruct the GPU to synthesize an image.

[0141] S616 . The GPU obtains the layer data from the memory 1 corresponding to the physical address 1 based on the mapping between the virtual address 2 and the physical address 1 , and synthesizes the image based on the layer data.

[0142] S617: The rendering service determines that the layer data needs to be synthesized through the DSS.

[0143] S618: The rendering service applies for a virtual address from the DSS, maps the virtual address 3 applied for by the DSS to the physical address 1, and sets the DSS flag.

[0144] In this embodiment of the present application, since the DSS flag bit is not set in S612, in other words, the virtual address has not yet been requested from the DSS and mapped to physical address 1, the rendering service requests the virtual address from the DSS. Optionally, in this embodiment of the present application, by setting the DSS flag bit, if the DSS is subsequently required to access memory 1 corresponding to physical address 1 again, the DSS flag bit can be used to detect that physical address 1 already has a DSS flag bit. In this way, the DSS can be instructed to access physical address 1 without having to request the virtual address from the DSS again and then map physical address 1.

[0145] It should be noted that if DSS has a flag bit, there is no need to apply for a virtual address from DSS.

[0146] S619: The rendering service instructs the DSS to synthesize the image.

[0147] In an embodiment of the present application, the rendering service may transmit an indication message 2 to the GPU, where the indication message 2 is used to instruct the DSS to synthesize an image.

[0148] S620 , DSS obtains layer data from memory 1 corresponding to physical address 1 based on the mapping between virtual address 3 and physical address 1 , and synthesizes an image based on the layer data.

[0149] In the embodiment of the present application, S614-S616 and S617-S620 are two parallel branches. In other words, the steps of S614-S616 may be executed or the steps of S617-S620 may be executed.

[0150] Optionally, the rendering service can choose whether to synthesize the image through the GPU or through the DSS according to certain business logic. In one possible implementation, the choice of synthesizing the image through the GPU or through the DSS can be made according to the power consumption and performance balance logic. Optionally, assuming that the performance of the GPU is better than the performance of the DSS and the power consumption of the GPU is greater than the power consumption of the DSS, at this time, if the rendering service detects that the electronic device is in power saving mode or the current power of the electronic device is lower than the power threshold, the rendering service can synthesize the image through the DSS; if the rendering service detects that the electronic device is in performance mode or not in power saving mode, or the current power of the electronic device is higher than or equal to the power threshold, the rendering service can choose to synthesize the image through the GPU. Optionally, the priority of synthesizing images for the GPU and DSS can also be set, and the image will be synthesized first through the memory with high priority, and when the memory with high priority is used up, the image will be synthesized through the memory with low priority.

[0151] It should be understood that the business logic of selecting a GPU or a CPU for image synthesis provided above is only an example and can be set as needed without limitation herein.

[0152] It should be noted that S611-S613 can be understood as the data production stage, and S614-S620 can be understood as the data consumption stage.

[0153] In another possible implementation, if the GPU is not involved in drawing the layer data, then in S612, the rendering service may request a virtual address from the CPU, allocate memory 1, map the virtual address 1 requested by the CPU to physical address 1, and set the CPU flag. Then, if the GPU needs to synthesize the image, before S615, the rendering service requests a virtual address from the GPU and maps the virtual address 2 requested by the GPU to physical address 1.

[0154] In another possible implementation, the GPU or DSS does not necessarily need to synthesize the image, but can also perform image enhancement processing on the layer data, which is not limited here.

[0155] exist Figure 6 In the illustrated embodiment, a case is described in which the CPU that runs the application and serves the rendering may be a different CPU from the CPU that writes the layer data to the physical address 1.

[0156] In the following embodiments, the CPU running the application and service rendering is the same CPU that writes layer data to memory 1 corresponding to physical address 1. In other words, the application and service rendering are software processes running on the CPU that writes layer data to memory 1 corresponding to physical address 1.

[0157] See also Figure 7 , Figure 7 This is a flow chart of another method for dynamically mapping shared physical memory to multiple device virtual addresses proposed in an embodiment of the present application. The method of this embodiment may include:

[0158] S711. The application requests graphics memory from the rendering service.

[0159] Among them, S711 can refer to the description of S611 and will not be repeated here.

[0160] S712: The rendering service applies for virtual address 1 of the CPU and virtual address 2 of the GPU, allocates memory 1, maps virtual address 1 and virtual address 2 applied by the GPU to physical address 1 of memory 1 respectively, and sets flags of the CPU and GPU.

[0161] In the embodiment of the present application, since the rendering service runs in the CPU that writes data, the rendering service can directly apply for virtual address 1.

[0162] It should be understood that it is also possible to apply for other services running in the CPU and then forward it to the rendering service for mapping between the virtual address and the physical address.

[0163] S713. The rendering service writes the layer data generated by the application into memory 1 based on the mapping between virtual address 1 and physical address 1.

[0164] In the embodiment of the present application, since the rendering service runs in the CPU that writes data, the rendering service can directly write the layer data generated by the application to the memory 1 corresponding to the physical address 1. Optionally, the layer data generated by the application can be transmitted to the rendering service and then written to the memory 1 corresponding to the physical address 1 by the rendering service.

[0165] It should be understood that the layer data generated by the application can also be written into the memory 1 corresponding to the physical address 1 through other services running in the CPU. In this case, the rendering service can transfer the layer data to other services, and then write the layer data into the memory 1 corresponding to the physical address 1 through other services.

[0166] S714: The rendering service determines that the layer data needs to be synthesized by the GPU.

[0167] Among them, S714 can refer to the description of S614 and will not be repeated here.

[0168] S715: The rendering service instructs the GPU to synthesize an image.

[0169] Among them, S715 can refer to the description of S615 and will not be repeated here.

[0170] S716 . The GPU obtains the layer data from the memory 1 corresponding to the physical address 1 based on the mapping between the virtual address 2 and the physical address 1 , and synthesizes the image based on the layer data.

[0171] Among them, S716 can refer to the description of S616 and will not be repeated here.

[0172] S717: The rendering service determines that the layer data needs to be synthesized through the DSS.

[0173] Among them, S717 can refer to the description of S617 and will not be repeated here.

[0174] S718: The rendering service applies for a virtual address from the DSS, maps the virtual address 3 requested by the DSS to the physical address 1, and sets the DSS flag.

[0175] Among them, S718 can refer to the description of S618 and will not be repeated here.

[0176] S719: The rendering service instructs the DSS to synthesize the image.

[0177] Among them, S719 can refer to the description of S619 and will not be repeated here.

[0178] S720 , DSS obtains layer data from memory 1 corresponding to physical address 1 based on the mapping between virtual address 3 and physical address 1 , and synthesizes an image based on the layer data.

[0179] Among them, S720 can refer to the description of S620 and will not be repeated here.

[0180] In the embodiment of the present application, in general, when an application in the application layer is started, the application can apply for graphics memory from the graphics memory management decision device (such as a rendering service). The graphics memory management decision device can first allocate the physical address of the producer device (such as a CPU) and the virtual address of the producer device, and set the flag bit of the producer device. The flag bit of the producer device can be used to mark the physical address allocated to the producer device, that is, the flag bit of the producer device can be used to mark the physical address allocated to the producer device as ready to be accessed by the producer device. Then, the graphics memory management decision device maps the virtual address of the producer device to the physical address of the memory allocated to the producer device. The producer device can use the mapping between the virtual address of the producer device and the physical address allocated to the producer device to write the produced data (such as image data) to the physical address of the memory allocated to the producer device. Then, the graphics memory management decision device selects the consumer device that needs to consume the data. If the consumer device does not set the flag bit, for example, if the DSS does not set the flag bit, then the consumer device's flag bit is set and the consumer device's virtual address is requested. The consumer device's flag bit indicates the physical address of the memory where the produced data is stored. In other words, the consumer device's flag bit can be used to mark the physical address storing the produced data as ready for access by the consumer device. The graphics memory management decision device then maps the consumer device's virtual address with the physical address storing the produced data. The consumer device can then utilize the mapping between the consumer device's virtual address and the physical address storing the produced data to retrieve the produced data from the memory corresponding to the physical address storing the produced data and consume the produced data.

[0181] In the embodiments of the present application, some scenarios for image processing using the CPU, GPU, and DSS may include but are not limited to:

[0182] Scenario 1: The CPU completes the drawing of the original layer and writes the drawing data into a buffer. The GPU or DSS then reads the drawing data from the buffer and uses it to synthesize the image to be displayed.

[0183] Scenario 2: The CPU completes the drawing of the original layer and writes the drawn data into a buffer. The GPU then reads the drawn data from the buffer and writes it into read-only memory (ROM) or an external device.

[0184] Scenario 3: The CPU completes the drawing of the original layer and writes the drawn data to a buffer. The GPU then reads the drawn data from the buffer and performs image enhancement processing, such as rendering, on the drawn data. The enhanced data is then written to the buffer. The DSS then reads the enhanced data from the buffer and uses it to synthesize the image to be displayed. Image enhancement processing may include, but is not limited to, one or more of contrast adjustment, sharpening, color enhancement, or filtering.

[0185] Scenario 4: The CPU and GPU complete the drawing of the original layer. Specifically, the CPU draws part of the original layer, while the GPU draws another part. The CPU and GPU each write the completed original layer to a buffer. The DSS then reads the drawn data from the buffer and uses it to synthesize the image to be displayed.

[0186] Scenario 5: The original layer drawing is completed by the CPU, and the drawing data is written to the buffer. The drawing data is then read from the buffer by the GPU and DSS, and the image to be displayed is synthesized by the GPU and DSS using the drawing data. In this case, optionally, the GPU can use a part of the drawing data to synthesize a part of the image, and the DSS can use another part of the drawing data to synthesize another part of the image, and then the partial image synthesized by the GPU and the other partial image synthesized by the DSS are spliced ​​to obtain a complete image. In another possible implementation, the GPU can use the drawing data to synthesize the image, and the DSS can use the drawing data to synthesize the image, and then the image synthesized by the GPU and the image synthesized by the DSS are fused to obtain a fused image.

[0187] The following embodiments respectively illustrate how to dynamically map shared physical memory to multiple device virtual addresses for the above situations.

[0188] For image processing corresponding to scenario 1, the rendering service can allocate memory for the CPU, request a virtual address for the CPU, and set a CPU flag. This CPU flag indicates the physical address of the memory allocated to the CPU. In other words, this CPU flag can be used to mark the physical address allocated to the CPU as ready for CPU access. The rendering service then maps the CPU's virtual address to the physical address. The CPU can then use the mapping between the CPU's virtual address and the physical address to write the image data (also known as drawn data or layer data) to the memory corresponding to the physical address. If the GPU is to be used for image synthesis, the rendering service sets a GPU flag. This GPU flag indicates the physical address where the image data is stored. In other words, this GPU flag can be used to mark the physical address where the image data is stored as ready for GPU access. The rendering service can also request a virtual address for the GPU and map the GPU's virtual address to the physical address where the image data is stored. The GPU can then use the mapping between the GPU's virtual address and the physical address where the image data is stored to retrieve the image data from the memory corresponding to the physical address for image synthesis. If an image needs to be synthesized using DSS, the rendering service sets the DSS flag, which indicates the physical address where the image data is stored. That is, the DSS flag can be used to mark the physical address where the image data is stored as ready for DSS access. The rendering service can also request the DSS virtual address and map the DSS virtual address to the physical address where the image data is stored. The DSS can then use the mapping between the DSS virtual address and the physical address where the image data is stored to retrieve the image data from the memory corresponding to the physical address of the image data for image synthesis.

[0189] If it is image processing corresponding to scenario 2, the rendering service can allocate memory for the CPU, apply for the virtual address of the CPU and set the flag of the CPU. Then, the rendering service maps the virtual address of the CPU to the physical address of the memory, and the CPU can use the mapping between the virtual address and the physical address of the CPU to write the image data to the physical address. Then, if the image data needs to be written to the read-only memory through the GPU, the rendering service sets the flag of the GPU, and the rendering service can apply for the virtual address of the GPU and map the virtual address of the GPU to the physical address where the image data is stored. The GPU can use the mapping between the virtual address of the GPU and the physical address of the image data to obtain the image data from the memory corresponding to the physical address of the image data and write it to the read-only memory.

[0190] If the image processing corresponds to scenario three, the rendering service can allocate memory for the CPU, request the CPU's virtual address, and set the CPU's flag. The rendering service then maps the CPU's virtual address to the physical address of the memory. The CPU can then use the mapping between the CPU's virtual address and the physical address to write the image data to the physical address. If image enhancement processing is required via the GPU, the rendering service sets the GPU's flag, requests the GPU's virtual address, and maps the GPU's virtual address to the physical address where the image data is stored. The GPU can then use the mapping between the GPU's virtual address and the physical address of the image data to retrieve the image data from the memory corresponding to the physical address of the image data for image enhancement processing. At this point, the GPU can write the enhanced image data to the physical address where the image data is stored, or to another physical address. If the enhanced image data is to be written to another physical address, the rendering service can allocate another memory for the GPU and set another flag for the GPU. This other flag indicates another physical address of the other memory. In other words, this other flag can be used to mark the other physical address allocated to the GPU as ready for access by the GPU. Furthermore, the rendering service can apply for another virtual address of the GPU and map the other virtual address of the GPU to another physical address. Then, the GPU can use the mapping between the other virtual address of the GPU and the other physical address to write the image data after image enhancement processing to the other physical address. Then, if it is necessary to synthesize the image through DSS, the rendering service sets the flag bit of DSS. The flag bit of DSS can be used to indicate the other physical address, that is, the flag bit of DSS can be used to mark that the other physical address is ready to be accessed by DSS. Furthermore, the rendering service can apply for the virtual address of DSS and map the virtual address of DSS to the other physical address. Then, DSS can use the mapping between the virtual address of DSS and the other physical address to obtain the image data after image enhancement processing from another memory corresponding to the other physical address for image synthesis.

[0191] It should be noted that in the image processing of scenario three, if the original image data needs to be retained, the image data after image enhancement processing can be written into another memory corresponding to another physical address; if the original image data does not need to be retained, the image data after image enhancement processing can be written into the memory corresponding to the physical address to replace the original image data previously written to the physical address.

[0192] For image processing corresponding to scenario 4, the rendering service can allocate memory for the CPU and GPU, requesting a CPU virtual address and a GPU virtual address. This can be understood as requesting virtual addresses for the CPU and GPU, and setting flags for the CPU and GPU. These flags indicate the physical addresses of the allocated memory for the CPU and GPU, marking the memory ready for access by the CPU or GPU. The rendering service then maps the CPU and GPU virtual addresses to physical addresses. The CPU can then use the mapping between the CPU's virtual address and the physical address to write image data to the memory corresponding to the physical address, and the GPU can also use the mapping between the GPU's virtual address and the physical address to write image data to the memory corresponding to the physical address. If the rendering service determines that the image data needs to be sent to the DSS for synthesis, it sets the DSS flag to indicate that the memory is ready for DSS access. The rendering service then requests the DSS's virtual address and maps the DSS's virtual address to the physical address. The DSS can then use the mapping between the DSS's virtual address and the physical address to retrieve image data from the memory corresponding to the physical address for synthesis.

[0193] If it is image processing corresponding to scenario five, the rendering service can allocate memory for the CPU, apply for the virtual address of the CPU and set the flag of the CPU. Then, the rendering service maps the virtual address of the CPU to the physical address, and the CPU can use the mapping between the virtual address of the CPU and the physical address to write the image data to the memory corresponding to the physical address. Then, apply for the virtual address of the CPU and the virtual address of the GPU, set the flags of the CPU and GPU, and map the virtual address of the CPU and the virtual address of the GPU to the physical address respectively. Then, the CPU can use the mapping between the virtual address applied by the CPU and the physical address to obtain the image data from the memory corresponding to the physical address for synthesis, and the GPU can use the mapping between the virtual address applied by the GPU and the physical address to obtain the image data from the memory corresponding to the physical address for synthesis.

[0194] In order to more intuitively illustrate the physical memory usage of the embodiment of the present application, the following embodiment provides diagrams to exemplarily illustrate the physical memory usage of the embodiment of the present application.

[0195] See also Figure 8 , Figure 8 This is a schematic diagram of physical memory usage provided by an embodiment of the present application. This embodiment of the present application is described as follows: the bit width of the CPU and GPU is 64 bits, and the bit width of the DSS is 32 bits. Figure 8As shown in the figure, the shared physical memory of CPU, GPU and DSS is 2G, while the shared physical memory of CPU and GPU is 6G. In other words, the shared physical memory that CPU and GPU can use is 8G. Figure 2 Compared with the physical memory usage of the CPU and GPU, the shared physical memory used by the CPU and GPU is improved, thereby improving the physical memory usage of the CPU and GPU, and thus improving the overall performance of the electronic device.

[0196] It should be noted that the physical memory shared by the CPU, GPU, and DSS is not limited to 2GB. It can be the maximum physical memory supported by the DSS. For example, if the DSS bit width is 32 bits, the maximum physical memory shared by the CPU, GPU, and DSS can not exceed 4GB. However, since the maximum physical memory supported by the CPU and GPU is 8GB, in this case, the shared physical memory shared by the CPU and GPU can be 4GB. In other words, the shared physical memory that the CPU and GPU can use can still be 8GB.

[0197] In this example, the physical memory usage can exceed 4G (the maximum bit width that can be used by the chip with the smallest bit width), and the virtual address of the chip with the smallest bit width does not exceed 4G.

[0198] In the embodiment of the present application, for example, for 16GB of physical memory, the shared memory utilization of the CPU and GPU is 8GB / 16GB = 50%, and for 32GB of physical memory, the shared memory utilization of the CPU and GPU is 8GB / 32GB = 25%. By comparing the utilization rates of the present application with those of the related art, the utilization rate of the shared memory in the embodiment of the present application is improved.

[0199] It should be understood that in the embodiments of the present application, the shared physical memory shared between multiple devices is limited by the maximum physical memory supported by the device with the smallest bit width. The shared physical memory between multiple devices with larger bit widths is further limited by the maximum physical memory supported by the device with the smallest bit width.

[0200] For example, assuming the CPU bit width is 128, the GPU bit width is 64, and the DSS bit width is 32 bits, the shared physical memory between the CPU, GPU, and DSS is limited by the maximum physical memory supported by the DSS (4GB). The shared physical memory between the CPU and GPU is also limited by the maximum physical memory supported by the GPU (8GB). In other words, the maximum shared memory between the CPU and GPU is 8GB, and the minimum is 4GB (the difference between the maximum physical memory supported by the GPU (8GB) and the maximum physical memory supported by the DSS (4GB).

[0201] For example, assuming the CPU bit width is 32 bits, the GPU bit width is 64 bits, and the DSS bit width is 128 bits, the shared physical memory between the CPU, GPU, and DSS is limited by the maximum physical memory supported by the CPU (4GB). The shared physical memory between the DSS and GPU is also limited by the maximum physical memory supported by the GPU (8GB). In other words, the maximum shared memory between the DSS and GPU is 8GB, and the minimum is 4GB (the difference between the maximum physical memory supported by the GPU (8GB) and the maximum physical memory supported by the CPU (4GB).

[0202] In the above embodiment, the scenario where the CPU, GPU, and DSS are one is described. In the following embodiment, the scenario where the CPU, GPU, and DSS are two or more is described.

[0203] In one possible implementation, assuming that there are multiple CPUs, including CPU3 and CPU4, when the CPU needs to produce data, the production data can be first written to the physical address by CPU3. When the address that CPU3 can use reaches the maximum physical memory supported by CPU3, the production data can be subsequently written to the physical address by CPU4. In this way, after one CPU reaches its usage limit, another CPU can be used, thereby improving the service life of some CPUs. Optionally, when the CPU needs to produce data, CPU3 and CPU4 can alternately write the production data to the physical address, for example, CPU3 writes the first batch of production data to physical address 3, CPU4 writes the second batch of production data to physical address 4, CPU3 writes the third batch of production data to physical address 5, CPU4 writes the fourth batch of production data to physical address 6, and so on, until the physical address used by at least one of CPU3 and CPU4 reaches the maximum physical memory supported by it. In this way, the overall service life of multiple CPUs can be improved.

[0204] In another possible implementation, assuming that there are multiple GPUs, the multiple GPUs include GPU1 and GPU2. If GPU production data is needed, the production data can be first written to the physical address through GPU1. When the address that GPU1 can use reaches the maximum physical memory supported by GPU1, the production data can be subsequently written to the physical address through GPU4. In this way, after one GPU reaches its usage limit, another GPU can be used, thereby improving the service life of some GPUs. Alternatively, GPU1 can be used specifically to write production data to the physical address, and GPU2 can be used specifically to read data from the physical address and process it. Optionally, GPU1 and GPU2 can process data alternately. For example, GPU1 writes the fifth batch of production data to physical address 7, and then GPU2 obtains the sixth batch of production data from physical address 8 for consumption. Then CPU1 obtains the seventh batch of production data from physical address 9 for consumption, and then GPU2 obtains the eighth batch of production data from physical address 10 for consumption. In this way, the overall service life of multiple GPUs can be improved.

[0205] It should be noted that the alternating processing of the GPU may be the alternating processing of production data, the alternating processing of consumption data, or the alternating processing of production data and consumption data, which is not limited here.

[0206] In another possible implementation, assuming there are multiple DSSs, including DSS1 and DSS2, if DSS consumption is required, data can first be obtained from a physical address through DSS1 for consumption. When the address that DSS1 can use reaches the maximum physical memory supported by DSS1, data consumption can continue through DSS1. In this way, after one DSS reaches its usage limit, another DSS can be used, thereby extending the service life of some DSSs. Optionally, DSS1 and DSS2 can process data alternately. For example, DSS1 obtains the ninth batch of production data from physical address 11 for consumption, and then obtains the tenth batch of production data from physical address 12 for consumption through DSS2. In this way, the overall service life of multiple DSSs can be extended.

[0207] It should be noted that, in the embodiments of the present application, how the CPU writes data to the physical address, how the GPU writes data to the physical address, how the GPU reads data from the physical address for processing, and how the DSS reads data from the physical address for processing can refer to the above embodiments. Figure 6 and Figure 7 The relevant description is not repeated here.

[0208] It should be understood that in the embodiments of the present application, the relationship between the devices performing data processing is mainly considered, and the size of the physical memory can meet the physical memory required by multiple devices to perform data processing by default.

[0209] In the above embodiments, the processing of graphic data is used as an example. Graphic data may include layer data. It should be understood that for other data processing scenarios, such as machine learning scenarios, including but not limited to the training or reasoning of artificial intelligence models, the training or reasoning of artificial intelligence models can also refer to the description of the embodiments of this application and will not be repeated here.

[0210] For example, the electronic device includes a data preprocessing chip (such as a CPU), a first training chip (such as a GPU), and a second training chip (such as an NPU). Among them, the data preprocessing chip can be used to preprocess the data, such as data cleaning, model selection and other preprocessing. Then, the model can be trained by the first training chip or the second training chip. Alternatively, the model is selected and pre-trained by the first training chip, and then the pre-trained model is continued to be trained by the second training chip. In this case, the data preprocessing chip acts as a producer device, the first training chip can act as a producer device or a consumer device, and the second training chip can act as a consumer device.

[0211] In this case, during the data production phase, virtual addresses can be requested from the data preprocessing chip and the first training chip respectively, and memory 2 can be allocated. The virtual address 4 requested by the preprocessing chip and the virtual address 5 requested by the first training chip can be mapped to the physical address 2 of memory 2. The data preprocessing chip can then write the selected model into the memory 2 corresponding to the physical address 2. Then, during the data consumption phase, if the first training chip is required for training, there is no need to request a virtual address from the first training chip. The first training chip can use the mapping between virtual address 5 and physical address 2 to obtain the model from the memory 2 corresponding to the physical address 2 for training. If the second training chip is required for training, a virtual address 6 can be requested from the second training chip, and virtual address 6 can be mapped to physical address 2. The second training chip can then obtain the model from the memory 2 corresponding to the physical address 2 based on the mapping between virtual address 6 and physical address 2 for training.

[0212] In addition, in the data production stage, after the first training chip selects a model and pre-trains the model, the first training chip uses the mapping between virtual address 5 and physical address 2 to write the pre-trained model into memory 2 corresponding to physical address 2. Then, in the data consumption stage, the second training chip can apply for virtual address 6 and map virtual address 6 to memory 2 corresponding to physical address 2. The second training chip can then obtain the pre-trained model from memory 2 corresponding to physical address 2 based on the mapping between virtual address 6 and physical address 2 to continue training.

[0213] It should be understood that, in addition to the above-mentioned example scenarios, this solution can also be applied to scenarios where a car head-up display is implemented through a GPU, such as displaying content on a car windshield.

[0214] In general, the shared memory allocation in the embodiment of the present application is divided into a production phase and a consumption phase, and each phase dynamically maps virtual addresses to related chips as needed.

[0215] See also Figure 9 , Figure 9 A flowchart of an address mapping method provided in an embodiment of the present application. Figure 9 The method shown can be applied to an electronic device, which includes a first application, a first device, a second device, a third device and a physical memory, wherein the bit width of the first device, the bit width of the second device and the bit width of the third device are not completely equal. For example, the bit width of the first device and the bit width of the second device may be greater than the bit width of the third device, or the bit width of the first device may be smaller than the bit width of the second device and the bit width of the third device, or the bit width of the second device may be smaller than the bit width of the first device and the bit width of the third device, and there is no limitation here. The first device, the second device and the third device may be chips, for example, the first device may include a CPU, the second device may include a GPU, and the third device may include a DSS. In an embodiment of the present application, the first device serves as a producer device, the second device may serve as both a producer device and a consumer device, and the third device may serve as a consumer device. As Figure 9 The method shown may include:

[0216] S901. A first application applies for memory.

[0217] The first application may be an application installed on an electronic device. In an embodiment of the present application, the first application may be an application that requires physical memory. For example, the first application may include, but is not limited to, a browser application and a camera application. For example, when a browser application or a camera application opens an application window, it requires memory to store graphic data.

[0218] For example, S901 may refer to the description of S611 and will not be elaborated here.

[0219] S902. Apply for a first virtual address of the first device and apply for a second virtual address of the second device, allocate a first memory in the physical memory, and map the first virtual address and the second virtual address to a first physical address of the first memory respectively. The first memory is at least used to store first data written by the first device.

[0220] A virtual address is a concept relative to a physical address. It is an address in the address space used to access memory. The main purpose of a virtual address is to provide an abstraction layer for memory management, enabling the operating system to effectively manage physical memory while maintaining memory isolation between processes. A physical address is a unique identifier for each byte in physical memory. It directly corresponds to an actual location in computer hardware (such as a memory chip).

[0221] In an embodiment of the present application, the first virtual address and the second virtual address are respectively mapped to the first physical address of the first memory, so that the first device and the second device can access the first physical address and then read and write data in the first memory. Optionally, the first device of this embodiment can act as a producer device to write the first data to the first memory. If the second device acts as a producer device, the second device can also write the second data of the second device to the first memory. If the second device acts as a consumer device, the second device can read the first data from the first memory for processing.

[0222] For example, if the first device includes a CPU, the first virtual address may be, for example, virtual address 1; if the second device includes a GPU, the second virtual address may be, for example, virtual address 2. The first memory may be, for example, memory 1, and the first physical address may be, for example, physical address 1. The third virtual address may be, for example, virtual address 3. S902 may refer to the description of S712 and will not be repeated here.

[0223] S903: Apply for a third virtual address of the third device, and map the third virtual address to the first physical address.

[0224] In an embodiment of the present application, when a third device is required to process data, that is, when the third device acts as a consumer device, the virtual address of the third device is requested and mapped to the first physical address, so that the third device can also access the first physical address, thereby obtaining data from the first memory for processing.

[0225] For example, S903 may refer to the description of S718 and will not be elaborated here.

[0226] S904: The third device obtains the first data from the first memory corresponding to the first physical address for processing.

[0227] In the embodiment of the present application, the processing that the third device is responsible for may be, for example, image synthesis, or model training or reasoning, etc., which is not limited here.

[0228] For example, S904 may refer to the description of S720 and will not be elaborated here.

[0229] It should be noted that since the first memory stores the first data, the third device can at least obtain the first data from the first memory for processing. If the second memory also stores the second data written by the second device, the third device can also obtain the second data from the first memory for processing.

[0230] In an embodiment of the present application, by applying for the first virtual address of the first device in the data production stage, allocating the first memory in the physical memory, and mapping the first virtual address to the first physical address of the first memory, the first device can write the first data to the first memory, and then apply for the second virtual address of the second device, and map the second virtual address to the first physical address of the first memory respectively. In this way, if the second device acts as a producer device, it can write the second data to the first memory, and if the second device acts as a consumer device, it can also obtain the first data from the first memory for processing. Then, when the third device acts as a consumer device, it can apply for the third virtual address of the third device and map the third virtual address to the first physical address. In this way, the third device can obtain the first data from the first memory corresponding to the first physical address for processing. Since the mapping of the virtual address is divided into multiple stages, such as the data production stage and the data consumption stage, the shared physical memory can be dynamically mapped to the virtual addresses of multiple chips, thereby improving the overall performance of the electronic device.

[0231] Exemplarily, the first data and the second data may be, for example, drawn layer data, and the corresponding data processing of the third device may be, for example, image synthesis processing.

[0232] In a possible implementation, after the first application applies for memory, the method may further include:

[0233] A first identifier of the first device and the second device is set, where the first identifier indicates that the device has been mapped to the first physical address.

[0234] In an embodiment of the present application, setting the first identifier of the first device indicates that the first device has been mapped to the first physical address. In other words, the virtual address of the first device applied for has been mapped to the first physical address. Setting the second identifier of the second device indicates that the second device has been mapped to the first physical address. In other words, the virtual address of the second device applied for has been mapped to the first physical address. Optionally, the identifier of the embodiment of the present application can also be called a flag bit.

[0235] It should be noted that if the device already has the first identifier, no mapping is required when the device needs to access the first physical address of the first memory.

[0236] For example, the embodiments of the present application can refer to the description of S612 and will not be elaborated here.

[0237] In the embodiment of the present application, after the virtual address of the device is mapped to the physical address, an identifier is set. In this way, it can be determined through the identifier whether the physical address is mapped to the virtual address of the device, so that data can be read and written to the memory corresponding to the physical address in a timely manner, thereby improving the efficiency of data processing. In addition, the number of virtual addresses of the device applied for can be reduced.

[0238] In another possible implementation, each time the device accesses the memory, it may be necessary to apply for the device's virtual address and map it to the physical address of the memory. This can reduce the resources required for setting the identifier.

[0239] In one possible implementation, applying for a third virtual address of a third device and mapping the third virtual address to the first physical address includes:

[0240] When it is determined that data is processed through a third device and the first identifier of the third device is not set, a third virtual address of the third device is applied for, and the third virtual address is mapped to the first physical address, and the first identifier of the third device is set, and the first identifier indicates that the device has been mapped to the first physical address.

[0241] For example, this embodiment can refer to the description of S614-S616, which will not be described in detail here.

[0242] In an embodiment of the present application, when it is determined that data is processed through a third device and the first identifier of the third device is not set, it means that the virtual address of the third device has not been mapped to the first physical address. Therefore, the third virtual address of the third device can be applied for and the third virtual address can be mapped to the first physical address. That is to say, when the third device acts as a consumer device, it is determined by the identifier whether it is necessary to apply for the virtual address of the third device and map it to the first physical address. This can improve the accuracy and efficiency of applying for the virtual address of the third device and mapping it.

[0243] In another possible implementation, when it is determined that data is processed through a third device, a virtual address of the third device is requested and mapped to the first virtual address, which can reduce the resources required for setting the identifier of the third device.

[0244] In a possible implementation, before applying for a third virtual address of the third device and mapping the third virtual address to the first physical address, the method further includes:

[0245] Obtain a first set and a second set, where the first set includes devices with a first identifier set, and the second set includes devices for processing data, where the devices with the first identifier set include the first device and the second device, and the devices for processing data include the third device;

[0246] If it is determined that a first difference set exists based on the first set and the second set, applying for a virtual device of a device in the first difference set, and mapping the virtual device of the device in the first difference set to a first physical address, the device in the first difference set is a device in the second set and the device in the first difference set is different from any device in the first set;

[0247] Applying for a third virtual address of a third device and mapping the third virtual address to the first physical address includes:

[0248] In a case where the first difference set includes the third device, a third virtual address of the third device is requested, and the third virtual address is mapped to the first physical address.

[0249] The first set may also be referred to as a set of marked device identifiers, and the second set may also be referred to as a set of consumer devices.

[0250] In this embodiment of the present application, the devices in the first difference set are identified consumer devices and are not marked with the first identifier. In other words, the devices in the first difference set need to apply for a virtual address and map it to the first physical address. If the first difference set includes a third device, it means that the third device needs to apply for a virtual address and map it to the first physical address. In this case, a third virtual address of the third device is applied for and mapped to the first physical address.

[0251] For example, the embodiments of this application can refer to Figure 5 The description of the embodiments will not be repeated here.

[0252] In a possible implementation, the electronic device further includes a first service, and the first application applies for memory, including:

[0253] The first application requests memory from the first service;

[0254] Applying for a first virtual address of a first device and applying for a second virtual address of a second device, allocating a first memory in a physical memory, and mapping the first virtual address and the second virtual address to a first physical address of the first memory, respectively, includes:

[0255] The first service applies for a first virtual address of the first device and a second virtual address of the second device, allocates a first memory in the physical memory, and maps the first virtual address and the second virtual address to a first physical address of the first memory respectively.

[0256] Exemplarily, the first service may be, for example, a rendering service, or other services running in the system, which is not limited here.

[0257] In an embodiment of the present application, the first application can apply for memory from the first service, and the first service can apply for the first virtual address of the first device and the second virtual address of the second device, allocate the first memory in the physical memory, and map the first virtual address and the second virtual address to the first physical address of the first memory respectively. In this way, the problems of inefficiency and high resource usage caused by the interaction of multiple services to achieve address mapping can be reduced, thereby improving the efficiency of address mapping and reducing resource usage.

[0258] In another possible implementation, virtual address application, memory allocation, and virtual-to-physical address mapping can be implemented by different services or modules. For example, a first service applies for virtual addresses, a second service allocates memory, and a third service maps virtual addresses to physical addresses, though this is not a limitation. This can reduce the problem of excessive service computing pressure caused by executing multiple processes through a single service.

[0259] In one possible implementation, applying for a third virtual address of a third device and mapping the third virtual address to the first physical address includes:

[0260] The first service applies for a third virtual address of the third device and maps the third virtual address to the first physical address.

[0261] In an embodiment of the present application, the first service can also apply for the virtual address of the third device and perform mapping. This can reduce the inefficiency and high resource usage caused by the interaction of multiple services to achieve address mapping, thereby improving the efficiency of address mapping and reducing resource usage.

[0262] In another possible implementation, the third virtual address of the third device can be applied for through other services, and the third virtual address can be mapped to the first physical address. In this way, the problem of excessive service computing pressure caused by executing multiple processes through one service can be reduced.

[0263] In a possible implementation, the first application applies for memory, including:

[0264] In response to opening a window on the electronic device, the first application requests a graphics memory, and the graphics memory is used to store graphics data required by the window.

[0265] In the embodiment of the present application, the window may be a window for starting the first application, or a newly opened window in the first application, which is not limited here.

[0266] In an embodiment of the present application, the first application requests graphics memory in response to opening a window on the electronic device, which can increase the number of windows opened on the electronic device.

[0267] In one possible implementation, the first memory is also used to store second data written by the second device, that is, the second device writes data to the first memory as a producer device.

[0268] In another possible implementation, applying for the second virtual address of the second device includes:

[0269] When it is determined that data is processed by the second device and the first identifier of the second device is not set, a second virtual address of the second device is applied for and the first identifier of the second device is set, where the first identifier indicates that the device has been mapped to the first physical address.

[0270] In the embodiment of the present application, the second device serves as a consumer device. If it is determined that the second device is required and the first identifier of the second device is not set, it means that the virtual address of the second device is not mapped to the first physical address, and the second device has difficulty accessing the first memory normally. Therefore, the second virtual address of the second device is requested, so that the second virtual address can be mapped to the first physical address, allowing the second device to access the first memory normally, and then obtain the first data in the first memory for processing.

[0271] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0272] It should also be understood that the various steps of the above embodiments may be coupled to each other, and this application does not limit this. Furthermore, the order of the sequence numbers of the above processes does not imply a specific order of execution. The execution order of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0273] The above describes in detail the address mapping method of the embodiment of the present application. Figure 10 , describes in detail the address mapping device of an embodiment of the present application.

[0274] Figure 10 1 shows a schematic block diagram of an address mapping device 1000 provided in an embodiment of the present application. The device 1000 includes a processor 1001, a transceiver 1002, and a memory 1003. The processor 1001, the transceiver 1002, and the memory 1003 communicate with each other via an internal connection path. The memory 1003 is used to store instructions, and the processor 1001 is used to execute the instructions stored in the memory 1003 to control the transceiver 1002 to send and / or receive signals.

[0275] It should be understood that the steps performed by the device 1000 in the embodiment of the present application can refer to the description of the above method embodiment and will not be repeated here.

[0276] In the embodiments of this application, Figure 10 The device 1000 may also be a chip or a chip system, such as a system on chip (SoC).

[0277] It should be understood that the device 1000 can be specifically the electronic device in the above-mentioned embodiment, and can be used to execute the various steps and / or processes corresponding to the electronic device in the above-mentioned method embodiment. Optionally, the memory 1003 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type. The processor 1001 can be used to execute instructions stored in the memory, and when the processor 1001 executes the instructions stored in the memory, the processor 1001 is used to execute the various steps and / or processes of the above-mentioned method embodiment. The transceiver 1002 may include a transmitter and a receiver. The transmitter can be used to implement the various steps and / or processes corresponding to the above-mentioned transceiver for performing the sending action, and the receiver can be used to implement the various steps and / or processes corresponding to the above-mentioned transceiver for performing the receiving action.

[0278] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0279] During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in a processor or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor executes the instructions in the memory, and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.

[0280] It should be understood that the apparatus 1000 herein may also be embodied in the form of functional modules. For example, the apparatus 1000 may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (e.g., a shared processor, a dedicated processor, or a group processor) and a memory for executing one or more software or firmware programs, combined logic circuits, and / or other suitable components that support the described functionality.

[0281] An embodiment of the present application further provides a computer-readable storage medium, which is used to store a computer program, and the computer program is used to implement the method shown in the above method embodiment.

[0282] An embodiment of the present application also provides a computer program product, which includes a computer program (also referred to as code or instructions). When the computer program runs on a computer, the computer can execute the method shown in the above method embodiment.

[0283] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0284] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0285] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0286] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present application.

[0287] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0288] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

Claims

1. An address mapping method, characterized in that: Applied to an electronic device, the electronic device includes a first application, a first device, a second device, a third device, and a physical memory, the bit width of the first device, the bit width of the second device, and the bit width of the third device are not completely equal, the method including: The first application applies for memory; Applying for a first virtual address of the first device and a second virtual address of the second device, allocating a first memory in the physical memory, and mapping the first virtual address and the second virtual address to a first physical address of the first memory, respectively, wherein the first memory is used to store at least first data written by the first device; Applying for a third virtual address of the third device, and mapping the third virtual address to the first physical address; The third device obtains the first data from the first memory corresponding to the first physical address for processing.

2. The method according to claim 1, characterized in that After the first application applies for memory, the method further includes: A first identifier of the first device and the second device is set, wherein the first identifier indicates that the device has been mapped to the first physical address.

3. The method according to claim 1 or 2, characterized in that The applying for a third virtual address of the third device and mapping the third virtual address to the first physical address includes: When it is determined that data is processed through the third device and the first identifier of the third device is not set, a third virtual address of the third device is applied for, the third virtual address is mapped to the first physical address, and the first identifier of the third device is set, and the first identifier indicates that the device has been mapped to the first physical address.

4. The method according to any one of claims 1 to 3, characterized in that Before applying for the third virtual address of the third device and mapping the third virtual address to the first physical address, the method further includes: Obtaining a first set and a second set, wherein the first set includes devices that have been set with the first identifier, and the second set includes devices used to process data, the devices that have been set with the first identifier include the first device and the second device, and the devices used to process data include the third device; If it is determined that a first difference set exists based on the first set and the second set, applying for a virtual device of a device in the first difference set, and mapping the virtual device of the device in the first difference set to the first physical address, the device in the first difference set is a device in the second set and the device in the first difference set is different from any device in the first set; The applying for a third virtual address of the third device and mapping the third virtual address to the first physical address includes: In a case where the first difference set includes the third device, a third virtual address of the third device is requested, and the third virtual address is mapped to the first physical address.

5. The method according to any one of claims 1 to 4, characterized in that The first device includes a central processing unit, the second device includes a graphics processing unit, and the third device includes a display subsystem.

6. The method according to any one of claims 1 to 5, characterized in that The electronic device further includes a first service, and the first application applying for memory includes: The first application requests memory from the first service; The applying for the first virtual address of the first device and the applying for the second virtual address of the second device, allocating the first memory in the physical memory, and mapping the first virtual address and the second virtual address to the first physical address of the first memory respectively, include: The first service applies for a first virtual address of the first device and a second virtual address of the second device, allocates a first memory in the physical memory, and maps the first virtual address and the second virtual address to a first physical address of the first memory respectively.

7. The method according to any one of claims 1 to 6, characterized in that The first application applies for memory, including: In response to opening a window on the electronic device, the first application requests a graphics memory, where the graphics memory is used to store graphics data required by the window.

8. The method according to any one of claims 1 to 7, characterized in that The first memory is further used to store second data written by the second device; or, The applying for the second virtual address of the second device includes: When it is determined that data is processed through the second device and the first identifier of the second device is not set, a second virtual address of the second device is applied for, and the first identifier of the second device is set, where the first identifier indicates that the device has been mapped to the first physical address.

9. An address mapping device, characterized in that: include: A processor is coupled to a memory, the memory stores computer-executable instructions, and the processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that Used to store a computer program, the computer program comprising instructions for implementing the method according to any one of claims 1 to 8.

11. A computer program product, comprising computer program code, characterized in that: When the computer program code is run on a computer, the computer is caused to implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Access of memory space exceeding address bus width

    CN107870870A

  • Method and device for communication between kernel mode and user mode and terminal

    CN108062253A

  • Data reading method and device and data storage method and device

    CN110502312A

  • Method for preloading memory page, electronic equipment and chip system

    CN115858046A

  • Memory access method, chip, electronic equipment and computer readable storage medium

    CN116136826A