Hardware calculation method and device, electronic equipment and storage medium
By scheduling at least two computing units in the computing module in parallel, the problem of low operating efficiency of hardware accelerator is solved, and more efficient data processing is achieved, suitable for fields such as image processing and artificial intelligence.
Patent Information
- Application Number
- CN202510412553.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-22
AI Technical Summary
The scheduling method of hardware accelerators in the prior art leads to low operational efficiency and low computing resource utilization rate.
By reading the pending data in memory and scheduling these units to perform hardware calculations in parallel according to the status of at least two computing units in the computing module, including releasing its resources when one computing unit is idle and calling another computing unit to process completed data, implementing parallel data processing of different computing units.
Improves computing efficiency, reduces data processing time, makes full use of hardware resources, and is suitable for scenarios that require a large amount of data calculations such as image processing and artificial intelligence.
Smart Images

Figure CN120353581A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a hardware computing method, apparatus, electronic device, and storage medium. Background Art
[0002] The modules for processing computing tasks in a computer include a vector processor and a hardware accelerator. The vector processor is a central processing unit that implements a direct operation instruction set for one-dimensional arrays (vectors), and the hardware accelerator is used to process tasks with a very large amount of computation. When processing tasks, the vector processor and the hardware accelerator can work together to process different types of computing tasks. However, in the related art, the scheduling method of the hardware accelerator results in low operating efficiency of the hardware accelerator. Summary of the Invention
[0003] This application aims to solve at least one of the technical problems in the related art to some extent.
[0004] To this end, this application proposes a method, apparatus, electronic device, and storage medium.
[0005] An embodiment of one aspect of this application proposes a hardware computing method, including:
[0006] Reading the data to be processed pre-written in the memory space;
[0007] According to the states of at least two computing units in the computing module, scheduling the at least two computing units to perform hardware computing on the data to be processed in parallel.
[0008] Optionally, the step of scheduling the at least two computing units to perform hardware computing on the data to be processed in parallel according to the states of at least two computing units in the computing module includes:
[0009] Scheduling the at least two computing units to perform different data processing in the hardware computing in parallel.
[0010] Optionally, the computing module includes a first computing unit and a second computing unit; the step of scheduling the at least two computing units to perform different data processing in the hardware computing in parallel includes any one of the following:
[0011] When the operating state of the first computing unit is computing completed, releasing the computing resources of the first computing unit, and calling the first computing unit to read the unprocessed data to be processed in the memory space and perform first data processing;
[0012] During the process of the first computing unit performing the first data processing, scheduling the second computing unit to perform second data processing on the data to be processed that the first computing unit has computed and completed.
[0013] Optionally, the scheduling for the at least two computing units to perform different data processing in parallel for hardware computing further includes:
[0014] When the operating state of the first computing unit is idle, the first computing unit is called to read the unprocessed data to be processed in the memory space and perform the first data processing.
[0015] Optionally, the scheduling for the second computing unit to perform second data processing on the data to be processed that has been computed by the first computing unit includes:
[0016] When the operating state of the second computing module is idle, the second computing unit is called to perform second data processing on the data to be processed that has been computed by the first computing unit.
[0017] Optionally, the method further includes:
[0018] When the second computing module has completed computing, the computing resources of the second computing unit are released, and the second computing unit is called to perform second data processing on the data to be processed that has been computed by the first computing unit.
[0019] Optionally, the method further includes:
[0020] Receive an instruction and determine the position of the data to be processed in the memory space according to the instruction, where the instruction is used to indicate performing hardware computing.
[0021] Optionally, the method further includes:
[0022] Output the processing result corresponding to the data to be processed;
[0023] Update the data in the indication area corresponding to the data to be processed.
[0024] Optionally, the method further includes:
[0025] Determine the computing progress of the data to be processed according to the indication areas corresponding to each group of data to be processed in the memory space.
[0026] Optionally, the method further includes:
[0027] In response to that the indication areas corresponding to each group of data to be processed all indicate completion of computing, perform decoding operation according to the decoding coefficient and the processing result.
[0028] Another embodiment of the present application proposes a hardware computing device, including:
[0029] A reading module, configured to read the data to be processed pre-written in the memory space;
[0030] A computing module, configured to schedule at least two computing units to perform hardware computing on the data to be processed in parallel according to the states of at least two computing units in the computing module.
[0031] Another embodiment of this application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the foregoing aspect is implemented.
[0032] Another embodiment of this application provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the foregoing aspect is implemented.
[0033] Another embodiment of this application provides a chip, which includes a processing circuit configured to execute the method described in the foregoing aspect.
[0034] Another embodiment of this application provides a computer program product, on which a computer program is stored. When the program is executed by a processor, the method described in the foregoing aspect is implemented.
[0035] The hardware computing method, device, electronic device, chip, and storage medium provided in this application read the data to be processed in the memory and schedule parallel computing according to the states of at least two computing units in the computing module. This design can make full use of hardware resources, improve computing efficiency, and reduce data processing time.
[0036] Additional aspects and advantages of this application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of this application. Description of the Drawings
[0037] The above and / or additional aspects and advantages of this application will become apparent and understandable from the following description of the embodiments in conjunction with the drawings, where:
[0038] Figure 1 is a schematic flowchart of a hardware computing method provided by an embodiment of this application;
[0039] Figure 2 is a schematic diagram of a data processing process provided by an embodiment of this application;
[0040] Figure 3 is a schematic diagram of a data processing process provided by an embodiment of this application;
[0041] Figure 4 is a schematic diagram of a data processing process provided by an embodiment of this application;
[0042] Figure 5 A structural schematic diagram of a hardware computing device provided by an embodiment of the present application;
[0043] Figure 6 A structural schematic diagram of an electronic device provided by an embodiment of the present application;
[0044] Figure 7 A structural schematic diagram of a chip proposed by an embodiment of the present application. Detailed implementation manners
[0045] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.
[0046] The modules for processing computing tasks in a computer include a vector processor and a hardware accelerator. The vector processor is a central processing unit that implements a direct operation instruction set for one-dimensional arrays (vectors), and the hardware accelerator is used to process tasks with a very large amount of computation.
[0047] The vector processor realizes the parallel execution of vector operations through a pipeline structure. It can perform the same operation on multiple data elements simultaneously within one clock cycle, greatly improving the computing efficiency. For example, in scientific computing, the vector processor can perform operations such as addition and multiplication on a set of data simultaneously, thereby accelerating the computing process. The hardware accelerator usually adopts a dedicated hardware design, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc., to achieve the efficient execution of specific algorithms or computing tasks. For example, a GPU is a common hardware accelerator, which accelerates tasks such as graphics rendering and matrix operations through a large number of parallel computing units.
[0048] In a computing system, according to the characteristics and requirements of the task, the task is reasonably divided between the vector processor and the hardware accelerator. For example, for scientific computing tasks that require a large amount of vector operations, the vector processor can be responsible for basic vector operations, while the hardware accelerator can accelerate specific complex computing parts therein, such as matrix multiplication.
[0049] Both can work simultaneously to jointly complete a complex computing task. The vector processor first performs preliminary processing on the data, then passes the result to the hardware accelerator for further optimization or specific operations, and finally integrates and outputs the final result. For example, in image processing, the vector processor can perform basic filtering operations on image pixels, and the hardware accelerator can accelerate more complex convolution operations.
[0050] The hardware calculation method, apparatus, electronic device, chip, and storage medium according to the embodiments of the present application will be described below with reference to the accompanying drawings.
[0051] Figure 1 It is a schematic flowchart of a hardware calculation provided by an embodiment of the present application.
[0052] As an implementation manner, the hardware calculation method of the embodiments of the present application can be configured in a hardware calculation device, and the hardware calculation device can be applied to any electronic device so that the electronic device can perform the hardware calculation function.
[0053] Among them, the electronic device can be any device with computing capabilities. For example, it can be a mobile terminal, and the mobile terminal can be a hardware device such as a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc. with various operating systems, touch screens, and / or display screens.
[0054] As another implementation manner, the hardware calculation method of the embodiments of the present application can also be executed by a chip with processing capabilities. The chip includes an Image Signal Processor (ISP), a Central Processing Unit (CPU), an Application-Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Field-Programmable Gate Array (FPGA), a System On A Chip (SOC), a Reduced Instruction Set Computer (RISC), etc., which will not be listed one by one here.
[0055] As Figure 1 shown, the method may include the following steps:
[0056] Step 101: Read the data to be processed pre-written in the memory space;
[0057] Step 102: According to the states of at least two computing units in the computing module, schedule the at least two computing units to perform hardware calculation on the data to be processed in parallel.
[0058] Figure 2 It is a schematic diagram of a data processing process provided by an embodiment of the present application. As Figure 2As shown in the figure, in the related art, a task requires processing of A, B, and module P. Among them, processing of A and B is performed by a vector processor (such as a central processing unit CPU), and module P is a hardware accelerator. The processing sequence is as follows: after performing processing of A, module P is called to process the output data, then processing of B is performed, and then module P is called again to perform operations on the output data.
[0059] After the processing of A or B is completed, in order to inform module P of the data, it is necessary to perform the interaction work between the software and hardware addresses. The process is as follows:
[0060] The data output by the processing of A or B is normalized;
[0061] The normalized data is placed in the memory space, and the data address is aligned;
[0062] Interface data transfer: Send signals to module P, including: interface register update signal, operation type signal of module P, size of each group of data, and starting address of Memory (memory). After receiving the signals, module P starts to read relevant information at the specified position and performs hardware calculations; assuming there are n groups of data, then n interface register update signals are sent;
[0063] During the process of module P performing accelerated calculations, the data in the indication area is read. The data in the indication area is used to indicate whether each group of data has been calculated. After the calculation is completed, the data in the indication area will become characteristic data (for example, when not processed, the data in the indication area is 0, and after the processing is completed, the data becomes 1).
[0064] In this process, each time module P is scheduled, it is serial. The processing process of each group of data is carried out in sequence, that is, after module P completes the accelerated calculation of a group of data, it can start the accelerated calculation of the next group of data. This results in module P not being in a working state all the time, with a high idle rate and low utilization rate of computing resources.
[0065] In this embodiment, the computing module is used to perform hardware calculations, and there are at least two computing units. The data to be processed is the data that needs to be processed by the computing module. During the process of processing multiple groups of data to be processed, at least two computing units are called to perform data processing in parallel to improve the efficiency of data processing. Optionally, the data to be processed is the data processed by the vector processor.
[0066] A basic framework of a hardware calculation method is proposed. By reading the data to be processed in the memory and scheduling parallel calculations according to the states of at least two computing units in the computing module. This design can make full use of hardware resources, improve the calculation efficiency, reduce the data processing time, and is applicable to scenarios that require a large amount of data calculation, such as the fields of image processing and artificial intelligence.
[0067] Optionally, step 102 schedules the at least two computing units to perform hardware calculations on the data to be processed in parallel according to the states of at least two computing units in the computing module, including:
[0068] Scheduling the at least two computing units to perform different data processing in the hardware calculation in parallel.
[0069] In this embodiment, the hardware calculation task is split into multiple levels of operations, and each level of budget is executed in different computing units. The calculation processes of different computing units are relatively independent, so that the computing units can process data in parallel.
[0070] The specific manner of parallel calculation is further refined, that is, scheduling at least two computing units to perform different data processing in the hardware calculation in parallel. This helps to better allocate tasks, avoid conflicts and redundancies between computing units, improve the parallelism and overall efficiency of data processing, and further enhance the performance of the system.
[0071] Optionally, the computing module includes a first computing unit and a second computing unit; the scheduling the at least two computing units to perform different data processing in the hardware calculation in parallel includes any one of the following:
[0072] When the operating state of the first computing unit is calculation completed, release the computing resources of the first computing unit, and call the first computing unit to read the unprocessed data to be processed in the memory space and perform first data processing;
[0073] During the process of the first computing unit performing first data processing, schedule the second computing unit to perform second data processing on the data to be processed that has been calculated and completed by the first computing unit.
[0074] In this embodiment, the first computing unit is used to execute the process of first data processing, and the second computing unit is used to execute the process of second data processing. The hardware calculation process of a set of data to be processed needs to first perform first data processing and then second data processing, so that the hardware calculation of this set of data to be processed is completed.
[0075] During the process of the first computing unit processing a set of data to be processed, since the second computing unit is relatively independent of the first computing unit, at this time, the second computing unit can simultaneously process another set of data to be processed that has undergone first data processing, without having to wait for the first computing unit to finish processing the data before performing data processing.
[0076] Similarly, during the process of the second computing unit processing a set of data to be processed, since the second computing unit is relatively independent of the first computing unit, the first computing unit can simultaneously process another set of unprocessed data to be processed at this time, without waiting for the second computing unit to finish processing the data before performing data processing.
[0077] It is clarified that the computing module includes a first computing unit and a second computing unit, and it is described in detail how to release its resources and call for new data processing when the first computing unit completes the calculation, while scheduling the second computing unit to perform subsequent processing on the completed data. Such a design realizes the dynamic allocation and efficient utilization of computing resources, avoids the idle and waste of resources, improves the flexibility and response speed of the system, and can better adapt to the changes in different data processing requirements.
[0078] Optionally, the scheduling of the at least two computing units to perform different data processing in parallel during hardware computing further includes:
[0079] When the operating state of the first computing unit is idle, call the first computing unit to read the unprocessed data to be processed in the memory space and perform the first data processing.
[0080] In this embodiment, the operating state of the first computing unit is monitored, and the operating states include: busy, idle, and calculation completed. It is in the busy state during the process of processing data, the state changes to the calculation completed state after the first data processing is completed, and changes to the idle state after releasing the computing resources.
[0081] It supplements the calling method when the first computing unit is idle, and further improves the scheduling mechanism of the computing unit. This enables the system to more timely utilize idle resources for data processing, reduces the task waiting time, improves the overall computing efficiency and throughput, and enhances the adaptability and stability of the system when processing burst data tasks.
[0082] Optionally, the scheduling of the second computing unit to perform second data processing on the data to be processed that has been calculated and completed by the first computing unit includes:
[0083] When the operating state of the second computing module is idle, call the second computing unit to perform second data processing on the data to be processed that has been calculated and completed by the first computing unit.
[0084] In this embodiment, the operating state of the second computing unit is monitored, and the operating states include: busy, idle, and calculation completed. It is in the busy state during the process of processing data, the state changes to the calculation completed state after the second data processing is completed, and changes to the idle state after releasing the computing resources.
[0085] The invocation condition for the second computing unit is restricted, i.e., it is invoked only when it is in an idle state. This helps to ensure that the second computing unit does not interfere with other ongoing tasks when performing new data processing tasks, guarantees the accuracy and reliability of data processing, optimizes the use of computing resources, avoids unnecessary resource occupation and performance loss, and further improves the stability and efficiency of the system.
[0086] Optionally, the method further includes:
[0087] In the case where the second computing module has completed the calculation, release the computing resources of the second computing unit, and invoke the second computing unit to perform second data processing on the data to be processed that has been completed by the first computing unit.
[0088] Figure 3 This is a schematic diagram of a data processing process provided by an embodiment of the present application. As Figure 3 described, in this embodiment, n groups of processed data to be processed are prepared in advance. First, the first group of data to be processed is input into the first computing unit for first data processing. After the first group of data to be processed is completed, the second computing unit is made to continue processing the first group of data to be processed. At the same time, the first computing unit reads the second group of data to be processed in the memory space and processes it. Assuming that the time required for the first computing unit and the second computing unit to perform calculations is both s clock cycles, then in an ideal situation, the time required to process n groups of data to be processed is about (s * n) / 2. The time required to schedule the computing module to process n groups of data to be processed using the serial scheduling method is s * n. The scheduling scheme in this embodiment shortens the time for hardware computing by nearly half.
[0089] It describes how to release the resources and re-invoke for new data processing after the second computing module has completed the calculation. This process realizes the recycling of computing resources, improves the resource utilization rate, enables the system to process more data tasks with limited hardware resources, enhances the continuous working ability and data processing ability of the system, and is of great significance for long-running computing tasks.
[0090] Optionally, the method further includes:
[0091] Receive an instruction and determine the position of the data to be processed in the memory space according to the instruction, where the instruction is used to indicate hardware computing.
[0092] In this embodiment, the instruction includes: an interface register update signal, an operation type signal, the size of each group of data to be processed, and the memory start address of the data to be processed. The interface register update signal is used to indicate the start of calculation. After receiving the instruction, the calculation module starts to read relevant information from the position of the memory start address and performs hardware calculation. Assuming there are n groups of data to be processed, there will be n interface register update signals.
[0093] The step of receiving an instruction and determining the position of the data to be processed in the memory space according to the instruction is added. This makes the calculation method more flexible and controllable. The user or the upper-level system can specify the data to be processed by sending an instruction, which facilitates the targeted calculation and processing of different data sets, improves the versatility and adaptability of the system, and can better meet diverse requirements.
[0094] Optionally, the method further includes:
[0095] Outputting the processing result corresponding to the data to be processed;
[0096] Updating the data in the indication area corresponding to the data to be processed.
[0097] In this embodiment, after each group of data to be processed undergoes the first data processing and the second data processing, in addition to generating the processed data, a shift value shift of the data will also be generated. The shift value is used to represent the address of the processed data in the memory space.
[0098] Each group of data to be processed has a corresponding indication area for indicating the calculation progress of the group of data to be processed.
[0099] The steps of outputting the processing result and updating the data in the indication area corresponding to the data to be processed are supplemented. This not only facilitates the user or the subsequent system to obtain the calculation result, but also enables the real-time monitoring and management of the calculation progress by updating the data in the indication area, helps to promptly discover and handle possible problems, improves the reliability and maintainability of the system, and at the same time provides accurate status information for subsequent data processing, which is conducive to the realization of the automation and intelligent management of the entire calculation process.
[0100] Optionally, the method further includes:
[0101] Determining the calculation progress of the data to be processed according to the indication areas corresponding to each group of data to be processed in the memory space.
[0102] In this embodiment, if the indication area changes, it indicates that the processing of this group of data to be processed has been completed.
[0103] In a possible embodiment, the initial data in the indication area is 0. After the first data processing and the second data processing are completed for the corresponding data to be processed, the data in the indication area changes to 1. By continuously reading the data in the indication area, when the data therein changes to 1, it can be determined that the data to be processed corresponding to the indication area has been processed.
[0104] By determining the calculation progress according to the indication areas corresponding to each group of data to be processed in the memory space, an intuitive and effective calculation progress management method is provided. This helps the user or the system to understand the processing situation of each group of data in real time, reasonably arrange subsequent tasks and resource allocation, improves the organization and coordination ability of the entire calculation process, and enhances the overall performance and working efficiency of the system.
[0105] Optionally, the method further includes:
[0106] In response to the indication areas corresponding to each group of data to be processed all indicating the completion of the calculation, perform a decoding operation according to the decoding coefficient and the processing result.
[0107] In this embodiment, after the hardware calculation is completed, a decoding operation needs to be performed. The pre-dependence data for the decode calculation is the output of the calculation module at the corresponding position and the decoding coefficient. Through the decoding operation, the data output by the calculation module is converted into a format suitable for subsequent processing and decoded.
[0108] After the indication areas corresponding to each group of data to be processed all indicate the completion of the calculation, perform a decoding operation according to the decoding coefficient and the processing result. This design combines the hardware calculation and the decoding operation, realizes the complete process of data processing, from the hardware calculation of the data to the final decoding output, improves the functional integrity and practicality of the system, and can meet more complex data processing requirements, such as application scenarios of quickly processing and decoding data in fields such as video coding and communication.
[0109] Figure 4 This is a schematic diagram of a data processing process provided by an embodiment of the present application. As Figure 4 described, in this embodiment, 8 groups of processing data are prepared and divided into 4×2 groups, and this grouping is used for the process of scheduling the decoding calculation. The data processing process is as follows:
[0110] The first group of data is first subjected to softening calculation in the vector processor, which is called the A processing process in this embodiment. The data output after the A processing is the data to be processed as described above;
[0111] Clear the calculation resources in the vector processor and call the vector processor to process the second group of data. At the same time, call the first calculation unit in the hardware calculation module to process the first group of data to be processed;
[0112] The first data processing of the first group of data to be processed in the first computing unit is completed, and the computing resources in the first computing unit are cleared. At this time, the A processing of the second group of data has been completed, and the first computing unit is called to process the second group of data to be processed; at the same time, the second computing unit in the hardware computing module is called to read the first group of data to be processed after the processing of the first computing unit and perform the second data processing;
[0113] After the second computing unit completes processing of the first group of data to be processed, the vector processor is called to process the third group of data, and the computing resources in the second computing unit are cleared. At this time, if the second group of data to be processed has been processed in the first computing unit, the second computing unit is called to perform the second data processing on the second group of data to be processed.
[0114] The processing flow of the subsequent vector processor and hardware computing module is similar.
[0115] After the A processing of the four groups of data is completed, the calculation of the decoding coefficient can be started, and then the indication area corresponding to the first four groups of data to be processed is read. If all the indication areas read have done signals, it means that the first data processing and the second data processing have been completed for these four groups of data to be processed. Next, the decoding calculation can be performed based on the decoding coefficient and the results of the processing of the four groups of data to be processed.
[0116] In order to implement the above embodiment, the embodiment of the present application also proposes a hardware computing device.
[0117] Figure 5 A schematic diagram of the structure of a hardware computing device provided in an embodiment of the present application.
[0118] like Figure 5 As shown, the device may include:
[0119] The reading module 510 is used to read the data to be processed that is pre-written in the memory space;
[0120] The computing module 520 is used to schedule the at least two computing units in the computing module to perform hardware computing on the to-be-processed data in parallel according to the states of the at least two computing units.
[0121] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment, and will not be repeated here.
[0122] In order to implement the above embodiments, the present application also proposes a non-temporary computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method described in the above method embodiments is implemented.
[0123] To implement the above embodiments, the present application also provides a computer program product, having a computer program stored thereon, and when the computer program is executed by a processor, the method described in the foregoing method embodiments is implemented.
[0124] To implement the above embodiments, the present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the foregoing method embodiments is implemented.
[0125] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0126] Refer to Figure 6 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0127] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0128] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0129] The power component 806 provides power for various components of the electronic device 800. The power component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0130] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0131] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0132] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0133] The sensor assembly 814 includes one or more sensors for providing an assessment of the status of various aspects of the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and a change in the temperature of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0134] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0135] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.
[0136] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 804 including instructions that can be executed by the processor 820 of the electronic device 800 to complete the above-described method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0137] To implement the above embodiments, the present application also proposes a chip, including: the chip includes a processing circuit configured to execute the method provided in the foregoing embodiments.
[0138] Figure 7 This is a schematic structural diagram of a chip proposed in an embodiment of the present application. Reference may be made to Figure 7 the schematic structural diagram of the chip 1100 shown, but not limited thereto.
[0139] The chip 1100 includes a processing circuit 1101, and the processing circuit 1101 is configured to execute any one of the above methods.
[0140] In some embodiments, the chip 1100 further includes one or more interface circuits 1102. Optionally, the interface circuit 1102 is connected to a memory 1103. The interface circuit 1102 can be used to receive signals from the memory 1103 or other devices, and the interface circuit 1102 can be used to send signals to the memory 1103 or other devices. For example, the interface circuit 1102 can read the instructions stored in the memory 1103 and send the instructions to the processing circuit 1101.
[0141] In some embodiments, the interface circuit 1102 executes at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 1101 executes other steps.
[0142] In some embodiments, terms such as interface circuit, interface, transceiver pin, transceiver, etc. can be replaced with each other.
[0143] In some embodiments, the chip 1100 further includes one or more memories 1103 for storing instructions. Optionally, all or part of the memories 1103 may be outside the chip 1100.
[0144] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0145] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of this application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0146] Any process or method description represented in a flowchart or described otherwise herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of this application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of this application belong.
[0147] The logic and / or steps represented in a flowchart or described otherwise herein, for example, may be considered as a sequenced list of executable instructions for implementing a logical function and may be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" may be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium may even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0148] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits with logic gate circuits for implementing logic functions on data signals, application specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0149] Those of ordinary skill in the art can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0150] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0151] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A hardware computing method, characterized in that, Including: Reading the data to be processed pre-written in the memory space; According to the states of at least two computing units in the computing module, scheduling the at least two computing units to perform hardware calculations on the data to be processed in parallel.
2. The method according to claim 1, characterized in that, The step of, according to the states of at least two computing units in the computing module, scheduling the at least two computing units to perform hardware calculations on the data to be processed in parallel includes: Scheduling the at least two computing units to perform different data processes in the hardware calculation in parallel.
3. The method according to claim 2, wherein The computing module includes a first computing unit and a second computing unit; the step of scheduling the at least two computing units to perform different data processes in the hardware calculation in parallel includes any one of the following: When the operating state of the first computing unit is calculation completed, releasing the computing resources of the first computing unit, and calling the first computing unit to read the unprocessed data to be processed in the memory space and perform a first data process; During the process of the first computing unit performing the first data process, scheduling the second computing unit to perform a second data process on the data to be processed that has been calculated and completed by the first computing unit.
4. The method according to claim 3, wherein The step of scheduling the at least two computing units to perform different data processes in the hardware calculation in parallel further includes: When the operating state of the first computing unit is idle, calling the first computing unit to read the unprocessed data to be processed in the memory space and perform the first data process.
5. The method according to claim 3, characterized in that, The step of scheduling the second computing unit to perform a second data process on the data to be processed that has been calculated and completed by the first computing unit includes: When the operating state of the second computing module is idle, calling the second computing unit to perform a second data process on the data to be processed that has been calculated and completed by the first computing unit.
6. The method according to claim 3, wherein The method further includes: When the second computing module is calculation completed, releasing the computing resources of the second computing unit, and calling the second computing unit to perform a second data process on the data to be processed that has been calculated and completed by the first computing unit.
7. The method according to claim 1, wherein The method further includes: Receiving an instruction and determining the position of the data to be processed in the memory space according to the instruction, where the instruction is used to indicate performing a hardware calculation.
8. The method according to claim 1, characterized in that The method further includes: Outputting the processing result corresponding to the data to be processed; Updating the data in the indication area corresponding to the data to be processed.
9. The method according to claim 8, wherein The method further includes: Determining the calculation progress of the data to be processed according to the indication areas corresponding to each group of data to be processed in the memory space.
10. The method according to claim 9, wherein, The method further includes: In response to the indication areas corresponding to each group of data to be processed all indicating that the calculation is completed, performing a decoding operation according to the decoding coefficient and the processing result.
11. A hardware computing device, characterized in that, Including: A reading module, configured to read the data to be processed pre-written in the memory space; A computing module, configured to schedule the at least two computing units to perform hardware calculations on the data to be processed in parallel according to the states of at least two computing units in the computing module.
12. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the method as described in any one of claims 1 - 10 is implemented.
13. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-10.
14. A chip, characterized in that, The chip includes a processing circuit configured to execute the method according to any one of claims 1-10.
15. A computer program product, characterized in that, It includes a computer program which, when executed by a processor, implements the method according to any one of claims 1-10.