Processor, vector aggregation method, computer equipment and storage medium

By combining butterfly networks and switching control circuits, high efficiency, accuracy and low resource consumption of vector convergence are achieved, solving the problem of high complexity of vector convergence operations in existing technologies and improving the performance of AI processors.

CN120973722APending Publication Date: 2025-11-18TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410612132.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-16
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

The vector convergence operation in existing AI processors is highly complex, resulting in large chip area, difficult placement and routing, and large input data fan-out of the multiplexer, leading to high resource consumption.

Method used

A butterfly network and a switch control circuit are used to achieve vector convergence through hierarchical switching nodes and a switch control module. The number of levels in the butterfly network has a logarithmic relationship with the length of the input vector. The switching nodes converge in pairs, and the switch control circuit precisely controls the switching nodes.

Benefits of technology

It reduces the layout area and wiring complexity of butterfly networks, saves resource consumption, and improves the accuracy of vector convergence and processor performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973722A_ABST
    Figure CN120973722A_ABST
Patent Text Reader

Abstract

The invention provides a processor, a vector convergence method, computer equipment and a storage medium, and belongs to the technical field of chips and semiconductors. The processor comprises a butterfly network and a switch control circuit; in any level of sub-network, the plurality of switching nodes in the sub-network are not connected; for the ith-level sub-network and the (i + 1) th-level sub-network, the jth switching node in the ith-level sub-network is respectively connected with the jth switching node and the nth switching node in the (i + 1) th-level sub-network, and the nth switching node in the ith-level sub-network is respectively connected with the nth switching node and the jth switching node in the (i + 1) th-level sub-network, the difference between n and j is 2i-1; and the switch control circuit is respectively connected with each switching node in the butterfly network. According to the processor, the wiring complexity is low, the layout cost is saved, and effective elements in the vector can be quickly converged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of chip and semiconductor technology, and in particular to a processor, vector convergence method, computer device and storage medium. Background Technology

[0002] Currently, AI (Artificial Intelligence) processors typically involve numerous vector position transformation operations. One such operation is vector convergence. Vector convergence involves selecting the valid data from two vectors and merging them into a single output vector. Because vector convergence has a large transformation space, data from any position can potentially be converged, resulting in high implementation complexity. How to achieve data convergence within vectors is a key research focus in this field.

[0003] In related technologies, data aggregation is achieved through multiplexers. For example, if the output vector length is m, which means the amount of data to be aggregated is m, then m multiplexers are deployed. The multiplexers are independent of each other.

[0004] However, as the parallelism of data processing in AI processors increases, a large number of multiplexers need to be deployed. Moreover, the multiplexers have numerous inputs, and as the vector length increases, the area of ​​the multiplexers becomes very large, which introduces a large number of connection lines. Each input data needs to drive m multiplexers, and the fan-out of the input data is large, resulting in a large chip area and greater difficulty in placement and routing. Summary of the Invention

[0005] This application provides a processor, a vector aggregation method, a computer device, and a storage medium, which not only saves layout costs but also facilitates the rapid aggregation of effective elements in a vector. The technical solution is as follows:

[0006] On one hand, a processor is provided, the processor comprising: a butterfly network and a switch control circuit;

[0007] The butterfly network includes multiple levels of sub-networks, and each level of sub-network includes multiple switching nodes;

[0008] In any sub-network, the multiple switching nodes in the sub-network are not connected;

[0009] For the i-th level subnetwork and the (i+1)-th level subnetwork, the j-th switching node in the i-th level subnetwork is connected to the j-th and n-th switching nodes in the (i+1)-th level subnetwork, respectively, and the n-th switching node in the i-th level subnetwork is connected to the n-th and j-th switching nodes in the (i+1)-th level subnetwork, respectively, where n and j differ by 2^(i-1).

[0010] The switch control circuit is connected to each switching node in the butterfly network respectively;

[0011] The switch control circuit is used to respond to a vector convergence command and generate control signals for each switching node in the butterfly network based on prompt information from multiple vectors, wherein the prompt information is used to indicate the valid elements in the multiple vectors;

[0012] Inputting multiple elements from the multiple vectors into the first-level subnetwork of the butterfly network;

[0013] For any switching node in the butterfly network, the switching node is used to swap the positions of two elements input to the switching node when the control signal is a first signal, and not to swap the positions of the two elements input to the switching node when the control signal is a second signal.

[0014] The last sub-network in the butterfly network outputs a target vector, which is composed of the valid elements of the plurality of vectors.

[0015] In some embodiments, the switch control circuit includes a plurality of switch control modules, and the switch control modules correspond one-to-one with the multi-level sub-network;

[0016] For any switch control module, the switch control module is connected to the switching node in the corresponding sub-network; the switch control module is used to respond to the vector convergence command, determine the state value of the multiple elements based on the prompt information of the multiple vectors, and generate the control signal of the corresponding switching node in the sub-network based on the state value of the multiple elements. The state value of each element is used to represent the number of valid elements between the element and the target element, and the target element is the least significant element among the multiple elements.

[0017] In some embodiments, the switch control module corresponding to the first-level sub-network in the butterfly network includes multiple LSB units, and the multiple LSB units correspond one-to-one with the switching nodes in the first-level sub-network.

[0018] For any LSB unit, the LSB unit is connected to the control terminal of the corresponding switching node, and the control terminal is used to receive the control signal generated by the LSB unit;

[0019] The LSB unit is configured to, based on the position of the switching node to which the LSB unit is connected in the primary sub-network, obtain a first element corresponding to the position of the switching node from a first type of elements among the plurality of elements, generate a first signal when the least significant bit of the state value of the first element is a first value, and generate a second signal when the least significant bit of the state value of the first element is a second value, wherein the first type of element is in an even position among the plurality of elements, and the state value of each element among the plurality of elements is a binary value.

[0020] In some embodiments, the switch control module corresponding to the primary sub-network in the butterfly network further includes a processing circuit;

[0021] The processing circuit is constructed based on the form of an XOR chain. The processing circuit includes multiple input terminals and multiple output terminals. The multiple input terminals are used to input the prompt information corresponding to each element in the multiple vectors respectively. The multiple output terminals are connected to the multiple LSB units respectively. The multiple output terminals are used to provide the least significant bit of the state value of the first type of element to the multiple LSB units.

[0022] In some embodiments, the switch control module corresponding to the non-first-level subnet in the butterfly network includes at least one LIR unit;

[0023] For any LIR unit, the LIR unit is connected to the control terminal of the corresponding switching node, and the control terminal is used to receive the control signal generated by the LIR unit;

[0024] The LIR unit is used to perform an offset operation on preset data based on the offset of the LIR unit to obtain target data, and generate a control signal based on the target data. The preset data is related to the level of the sub-network corresponding to the LIR unit, and the offset is the number of times the offset operation is performed.

[0025] In any offset operation, the LIR unit is used to extract the value of the highest bit in the preset data, shift the remaining value in the preset data one bit to the left, invert the value of the highest bit, and place the inverted value in the lowest bit of the preset data.

[0026] In some embodiments, for any LIR unit, the LIR unit is used to obtain a second element corresponding to the position of the switching node from a second type of elements among the plurality of elements based on the position of the switching node in the sub-network to which the LIR unit is connected, and to determine the offset based on the state value of the second element, wherein the second type of element is in an odd position among the plurality of elements.

[0027] In some embodiments, for any LIR unit, the LIR unit is used to obtain a target number of low-order values ​​from the state value of the second element corresponding to the LIR unit, and use the values ​​as the offset, wherein the target number is related to the level of the sub-network corresponding to the LIR unit.

[0028] On the other hand, a vector convergence method is provided, the method being executed by a processor, the processor being the processor according to any one of claims 1 to 7;

[0029] The method includes:

[0030] In response to a vector convergence command, control signals for each switching node in the butterfly network of the processor are generated based on prompt information from multiple vectors, wherein the prompt information is used to indicate the valid elements in the multiple vectors;

[0031] The plurality of vectors are input into the butterfly network, and based on the control signals of each switching node in the butterfly network, the flow of multiple elements in the plurality of vectors in the butterfly network is controlled.

[0032] For any switching node, if the control signal of the switching node is the first signal, the two elements input in the switching node are swapped; if the control signal is the second signal, the two elements input in the switching node are not swapped.

[0033] Output a target vector, which is composed of the valid elements of the plurality of vectors.

[0034] On the other hand, a computer device is provided, the computer device including a processor and a memory, the processor being the processor described in any of the above embodiments, the memory being used to store at least one computer program, the at least one computer program being loaded and executed by the processor to implement the vector convergence method in the embodiments of this application.

[0035] On the other hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, the at least one computer program being loaded and executed by a processor to implement the vector convergence method as described in the embodiments of this application.

[0036] On the other hand, a computer program product is provided, including a computer program stored in a computer-readable storage medium, a processor of a computer device reading the computer program from the computer-readable storage medium, and the processor executing the computer program, causing the computer device to perform the vector convergence method provided in the above aspects or various alternative implementations of the above aspects, wherein the processor is the processor described in any of the above embodiments.

[0037] This application provides a processor that uses a butterfly network to aggregate valid elements from multiple vectors. Since the butterfly network performs data selection hierarchically, with the number of levels and the length of the input vector having a logarithmic relationship, the layout area of ​​the butterfly network is significantly reduced. Furthermore, each sub-network uses a pairwise aggregation method with swapping nodes to exchange the positions of elements in multiple vectors, resulting in lower wiring complexity. This not only saves overall resource consumption and layout costs but also facilitates faster acquisition of valid elements from the vectors. Moreover, by connecting a switch control circuit to each swapping node in the butterfly network, precise control of each swapping node is achieved, thereby more accurately controlling the flow direction of elements in the vectors and more accurately acquiring valid elements, thus improving processor performance. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 This is a schematic diagram of a processor provided in an embodiment of this application;

[0040] Figure 2 This is a schematic diagram of a switching node provided according to an embodiment of this application;

[0041] Figure 3 This is a schematic diagram of another processor provided according to an embodiment of this application;

[0042] Figure 4 This is a schematic diagram of another processor provided according to an embodiment of this application;

[0043] Figure 5 This is a schematic diagram of a processing circuit provided according to an embodiment of this application;

[0044] Figure 6 This is a schematic diagram of an offset operation provided according to an embodiment of this application;

[0045] Figure 7 This is a schematic diagram of another processing circuit provided according to an embodiment of this application;

[0046] Figure 8 This is a schematic diagram of the implementation environment of a vector convergence method provided in an embodiment of this application;

[0047] Figure 9 This is a flowchart of a vector convergence method provided according to an embodiment of this application;

[0048] Figure 10 This is a structural block diagram of a terminal provided according to an embodiment of this application;

[0049] Figure 11 This is a schematic diagram of the structure of a server according to an embodiment of this application. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0051] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor are there any restrictions on quantity or execution order.

[0052] In this application, the term "at least one" means one or more, and "multiple" means two or more.

[0053] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the vectors involved in this application were all obtained with full authorization.

[0054] For ease of understanding, the terms used in this application are explained below.

[0055] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0056] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0057] This application provides a processor, which is an AI processor developed based on artificial intelligence technology and can be applied to artificial intelligence chips. The processor includes a butterfly network and a switching control circuit.

[0058] First, let's introduce the butterfly network in the processor. A butterfly network consists of multiple levels of subnetworks. Each level of subnetwork includes multiple swapping nodes. The total number of levels in the butterfly network is related to the length of the multiple input vectors. The length of a vector refers to the number of elements in the vector. The total number of levels of the subnetwork is N = log2(m), where m is the number of elements in the multiple vectors. For example, m = 8, so the butterfly network has 3 levels, and each level of subnetwork has m / 2 swapping nodes.

[0059] In any subnetwork, multiple switching nodes in that subnetwork are not connected.

[0060] In a butterfly network, any two adjacent subnetworks are connected, allowing input data to flow sequentially through the subnetworks and ultimately reach the desired location. Within any two adjacent subnetworks, nodes at different subnetwork levels are connected. The connection pattern resembles that of a butterfly. The following explanation uses the i-th and (i+1)-th subnetworks as examples.

[0061] For the i-th level subnetwork and the (i+1)-th level subnetwork, the j-th switching node in the i-th level subnetwork is connected to the j-th and n-th switching nodes in the (i+1)-th level subnetwork, respectively. The n-th switching node in the i-th level subnetwork is connected to the n-th and j-th switching nodes in the (i+1)-th level subnetwork, respectively. Here, n and j differ by 2. i-1 In j / 2 i The remainder is no greater than 2 i-1 In the case of n = j + 2 i-1 ; in j / 2 i The remainder is greater than 2 i-1 Or j / 2 i In the case of no remainder, n = j - 2 i-1 .

[0062] For example, Figure 1 This is a schematic diagram of a processor provided in an embodiment of this application. See also... Figure 1 The processor includes a butterfly network 10 and a switch control circuit 11. The number of switching nodes in each sub-network of the butterfly network 10 is the same, four switching nodes in each level. The switching nodes in the first-level sub-network are S00, S01, S02, and S03; the switching nodes in the second-level sub-network are S10, S11, S12, and S13; and the switching nodes in the third-level sub-network are S20, S21, S22, and S23.

[0063] Let's take the first-level subnetwork and the second-level subnetwork as examples to introduce the connection method of the switching nodes in the butterfly network. i=1.

[0064] If j = 1, then j / 2 i The remainder is 1, and 1 is not greater than 2. 1-1 Accordingly, the first exchange node in the first-level subnetwork is connected to the first exchange node and the second exchange node in the second-level subnetwork (n = 1 + 2). 1-1 =1+2 0 =1+1=2) Swap the nodes and connect them respectively. That is, Figure 1 The switching node S00 is connected to switching nodes S10 and S11 respectively.

[0065] If j = 2, then j / 2 i There is no remainder. Correspondingly, the second exchange node in the first-level subnetwork and the second exchange node in the second-level subnetwork are the same as the first (n = 2 - 2) exchange node. 1-1 =2-2 0 =2-1=1) swap nodes are connected respectively. That is, Figure 1 The exchange node S01 is connected to exchange nodes S11 and S10 respectively.

[0066] If j = 3, then j / 2 i The remainder is 1, and 1 is not greater than 2. 1-1 Correspondingly, the third exchange node in the first-level subnetwork is connected to the third exchange node and the fourth exchange node in the second-level subnetwork (n = 3 + 2). 1-1 =3+2 0 =3 + 1 = 4) swap nodes are connected respectively. That is, Figure 1 The exchange node S02 is connected to exchange nodes S12 and S13 respectively.

[0067] If j = 4, then j / 2 i There is no remainder. Correspondingly, the 4th exchange node in the first-level subnetwork and the 4th exchange node and the 3rd (n = 4 - 2) in the second-level subnetwork... 1-1 =4-2 0 =4-1=3) swap nodes are connected respectively. That is, Figure 1 The exchange node S03 is connected to exchange nodes S13 and S12 respectively.

[0068] The butterfly network inputs multiple elements from multiple vectors into its first-level subnetwork. These multiple vectors refer to the vectors to be converged. The processor can input the elements from these multiple vectors into the various exchange nodes in the first-level subnetwork of the butterfly network, controlling the flow of these elements within the butterfly network so that the valid elements to be converged flow to the designated positions. Correspondingly, the last-level subnetwork of the butterfly network outputs the target vector. The target vector consists of the valid elements from the multiple vectors.

[0069] Before inputting multiple vectors into the butterfly network, the switching control circuit can generate control signals for each switching node in the butterfly network based on the prompt information of the multiple vectors, so that the butterfly network can control the flow direction of elements in the multiple vectors according to the control signals of each switching node.

[0070] The following describes the switching control circuit in the processor. The switching control circuit is connected to each switching node in the butterfly network.

[0071] A switching control circuit, in response to a vector convergence command, generates control signals for each switching node in the butterfly network based on prompts from multiple vectors. The prompts indicate the valid elements within the multiple vectors.

[0072] The prompt information for the vector can be a corresponding Boolean vector (Vector Boolean, VB), and this embodiment of the application does not impose any restrictions on this. Each Boolean value in the Boolean vector corresponds one-to-one with an element in the vector. Each Boolean value indicates whether the corresponding element is a valid element. A valid element is the element to be aggregated. For example, a Boolean value of 1 indicates that the corresponding element is a valid element; a Boolean value of 0 indicates that the corresponding element is not a valid element, i.e., an invalid element.

[0073] For example, see [link to relevant documentation] Figure 1 The system consists of multiple vectors, including a first vector and a second vector. The first vector is [hg fe]. The prompt information for the first vector is [1 0 0 1]. Therefore, the valid elements in the first vector are h and e. The second vector is [dcba]. The prompt information for the second vector is [0 1 1 0]. Therefore, the valid elements in the second vector are c and b. The processor can combine the prompt information of the first vector and the prompt information of the second vector into a single vector [1 0 0 10 1 1 0], and input it into the switch control circuit. This allows the switch control circuit to generate control signals for each switching node in the butterfly network based on this vector (prompt information). Then, the processor inputs the elements in the vectors into each switching node of the primary subnetwork in the butterfly network, according to the order of the elements. Each switching node inputs two elements. Then, based on the control signals of each switching node in the butterfly network, the processor controls the flow of the eight elements in the first and second vectors within the butterfly network.

[0074] During the flow of elements in the butterfly network, for any exchange node in the butterfly network, the exchange node is used to swap the positions of the two elements input to the exchange node when the control signal is the first signal. When the control signal is the second signal, the two elements input to the exchange node are not swapped.

[0075] For any given switching node, it includes two inputs and two outputs: a first input, a second input, a first output, and a second output. The two inputs and two outputs are aligned one-to-one. That is, the first input is aligned with the first output, and the second input is aligned with the second output. Position swapping means that data input from the first input is output from the second output, and data input from the second input is output from the first output. Not swapping means that data input from the first input is output from the first output, and data input from the second input is output from the second output.

[0076] For example, Figure 2 This is a schematic diagram of a switching node according to an embodiment of this application. See also... Figure 2The switching node includes a first input terminal A1, a second input terminal B1, a first output terminal A2, and a second output terminal B2. The element input to the first input terminal A1 is 'a'; the element input to the second input terminal B1 is 'b'. When the control signal is the first signal, the switching node swaps the positions of elements 'a' and 'b'. That is, the first output terminal A2 outputs element 'b'; the second output terminal B2 outputs element 'a'. (See [reference needed]). Figure 2 (1) When the control signal is the second signal, the switching node does not swap the positions of elements a and b. That is, the first output terminal A2 of the switching node outputs element b; the second output terminal B2 outputs element a. See also: Figure 2 (2) In this context, the control signal is an electrical signal. The first signal can be 0 (low voltage signal), and the second signal can be 1 (high voltage signal). This application does not impose any restrictions on this. The processor controls the flow of elements in multiple vectors between multiple switching nodes in the butterfly network based on the control signals of each switching node in the butterfly network.

[0077] This application provides a processor that uses a butterfly network to aggregate valid elements from multiple vectors. Since the butterfly network performs data selection hierarchically, with the number of levels and the length of the input vector having a logarithmic relationship, the layout area of ​​the butterfly network is significantly reduced. Furthermore, each sub-network uses a pairwise aggregation method with swapping nodes to exchange the positions of elements in multiple vectors, resulting in lower wiring complexity. This not only saves overall resource consumption and layout costs but also facilitates faster acquisition of valid elements from the vectors. Moreover, by connecting a switch control circuit to each swapping node in the butterfly network, precise control of each swapping node is achieved, thereby more accurately controlling the flow direction of elements in the vectors and more accurately acquiring valid elements, thus improving processor performance.

[0078] To more clearly describe how the switch control circuit controls the switching nodes in the butterfly network, the switch control circuit will be further introduced below.

[0079] In some embodiments, the switch control circuit includes multiple switch control modules. Each switch control module corresponds one-to-one with a multi-level sub-network. That is, the switch control circuit controls the multi-level sub-networks separately through each switch control module. Each switch control module is independent of the others. This approach facilitates pipelined design, allowing for flexible selection of the required switch control modules based on the layout of the butterfly network, thus enabling flexible control of the switching nodes in the corresponding sub-networks.

[0080] For any given switch control module, it is connected to a switching node in the corresponding sub-network. The switch control module, in response to a vector convergence command, determines the state values ​​of multiple elements based on prompts from multiple vectors; then, based on the state values ​​of these multiple elements, it generates control signals for the corresponding switching node in the sub-network. The state value of each element represents the number of valid elements between that element and the target element. The target element is the least significant element among the multiple elements.

[0081] The prompts from multiple vectors are Boolean vectors. Each element corresponds to a Boolean value. For any given element, the switch control module accumulates the Boolean values ​​of all elements between the target element and that element to obtain the state value of that element.

[0082] For example, Figure 3 This is a schematic diagram of another processor provided according to an embodiment of this application. See also... Figure 3 The processor includes a butterfly network 30 and a switch control circuit 31. The input elements in the butterfly network 30 are hgfedcba. The Boolean values ​​corresponding to the elements hgfedcba are 1 0 0 1 0 1 1 0. Element a is the target element. Taking element e as an example, the state value of element e = 0 + 1 + 1 + 0 + 1 = 3, which is 010 in binary. The number of bits in the binary representation of the state value is equal to log2(m), where m is the number of elements in the multiple vectors. Based on the above method, the state values ​​of hgfedcba are calculated as 100, 011, 011, 011, 010, 010, 001, and 000. The switch control modules 310, 311, and 312 in the switch control circuit 31 generate control signals for the corresponding switching nodes in the sub-network based on the state values ​​of the multiple elements.

[0083] The solution provided in this application embodiment controls whether the switching node performs position swapping on the input elements based on the number of valid elements to be converged between each element and the target element. This enables accurate control of the flow direction of valid elements, bringing them together and improving the accuracy of vector convergence.

[0084] In this embodiment, the processor can employ different switch control modules with different structures to control the switching nodes in different levels of sub-networks. That is, the processor uses different methods to generate control signals for the switching nodes in each level of the sub-network. The switch control modules corresponding to the first-level sub-network and the non-first-level sub-networks in the butterfly network are described below.

[0085] In this embodiment of the application, the switch control module corresponding to the first-level subnetwork in the butterfly network includes multiple LSB (Least Significant Bit) units. Each LSB unit corresponds one-to-one with a switching node in the first-level subnetwork.

[0086] For any LSB unit, the LSB unit is connected to the control terminal of the corresponding switching node. The control terminal is used to receive control signals generated by the LSB unit.

[0087] For any LSB unit, based on the position of the switching node connected to that LSB unit in the head-level subnetwork, it retrieves the first element corresponding to the position of the switching node from a first type of elements among multiple elements. If the least significant bit of the first element's state value is a first value, a first signal is generated. If the least significant bit of the first element's state value is a second value, a second signal is generated. The first type of elements are located at even-numbered positions among the multiple elements. The state value of each element among the multiple elements is a binary value.

[0088] For example, Figure 4 This is a schematic diagram of another processor provided according to an embodiment of this application. See also... Figure 4 LSB00 unit is connected to switching node S00; LSB01 unit is connected to switching node S01; LSB02 unit is connected to switching node S02; LSB03 unit is connected to switching node S03.

[0089] The first type of elements among multiple elements are the elements located at positions 0, 2, 4, and 6, namely elements a, c, e, and g, respectively. The LSB00 unit can obtain the first element corresponding to the position of the switching node S00 connected to the LSB00 in the primary subnetwork. Position correspondence means that the order of the switching node in the subnetwork is equal to the order of the first element in the first type of elements. For example, if the switching node S00 is the first switching node in the primary subnetwork, then the first element corresponding to the switching node S00 is the first first type of element, which is element a. The state value of element a is 000. The least significant bit of the state value of element a is the first numerical value (0). Accordingly, the LSB00 unit generates the first signal corresponding to the first numerical value (0). Following the above method, the LSB01 unit generates the first signal corresponding to the first numerical value (0), the LSB02 unit generates the second signal corresponding to the second numerical value (1), and the LSB03 unit generates the second signal corresponding to the second numerical value (1). Based on this, in the primary sub-network, swapping nodes S00 and S01 will swap the positions of the input elements; swapping nodes S02 and S03 will not swap the positions of the input elements.

[0090] In some embodiments, the switch control module corresponding to the primary subnetwork in the butterfly network further includes a processing circuit. See also... Figure 4 Each sub-network's corresponding switch control module includes a processing circuit. This processing circuit calculates the state value of each element. It is connected to the switch control unit and provides the switch control unit with the state values ​​of each element.

[0091] The processing circuit corresponding to the first-level sub-network is constructed based on an XOR chain structure. The processing circuit includes multiple input terminals and multiple output terminals. The multiple input terminals are used to input the prompt information corresponding to each element in multiple vectors. The multiple output terminals are connected to multiple LSB units respectively. The multiple output terminals are used to provide the least significant bit of the state value of the first type of element to the multiple LSB units.

[0092] For example, Figure 5 This is a schematic diagram of a processing circuit according to an embodiment of this application. See also... Figure 5 Since the switching nodes in the primary subnetwork are controlled only by the least significant bit of the state value of the first type of element, there is no need to perform a full-width addition calculation on the state values. Only an XOR operation on the Boolean values ​​of each element is required. Here, an XOR chain is used for calculation, and the result at even-numbered positions is extracted as the output. This method is relatively simple to calculate and has low latency.

[0093] In this embodiment, the switch control module corresponding to a non-first-level subnet in the butterfly network includes at least one LIR (Left Invert Rotate) unit. The number of LIR units in the switch control module corresponding to each subnet is equal to m / 2. i m is the number of elements in the multiple vectors. The number of switching nodes controlled by the LIR unit in each sub-network level is 2. i-1 .

[0094] For any LIR unit, the LIR unit is connected to the control terminal of the corresponding switching node. The control terminal is used to receive control signals generated by the LIR unit.

[0095] For example, see continue. Figure 4The second-level subnetwork's switch control module contains two LIR units: LIR10 and LIR11. Each LIR unit controls two switching nodes. LIR10 is connected to switching nodes S10 and S11 respectively; LIR11 is connected to switching nodes S12 and S13 respectively. The third-level subnetwork's switch control module contains one LIR unit: LIR20. LIR20 controls four switching nodes: S20, S21, S22, and S23 respectively.

[0096] For any given LIR unit, the LIR unit performs an offset operation on preset data based on the shift amount of that LIR unit to obtain target data, and then generates a control signal based on the target data. The preset data is related to the level of the sub-network corresponding to the LIR unit. The shift amount is the number of times the offset operation is performed.

[0097] For any offset operation, the LIR unit is used to extract the value of the highest bit in the preset data, shift the remaining value in the preset data one bit to the left, invert the value of the highest bit, and place the inverted value in the lowest bit of the preset data.

[0098] In this embodiment, both LIR10 and LIR11 are 2-bit offset operations. That is, the preset data is 00. For example, Figure 6 This is a schematic diagram of an offset operation according to an embodiment of this application. The preset data (initial data) is "00". During the first offset, the LIR unit extracts the value "0" of the highest bit (leftmost bit) of "00", shifts the remaining "0" one bit to the left, inverts the highest bit "0" to obtain "1", and places "1" in the lowest bit of the preset data, thus obtaining "01". Then, the LIR unit can perform a second offset based on "01" to obtain "11"; perform a third offset based on "11" to obtain "10"; perform a fourth offset based on "11" to obtain "00", and so on. The number of offsets depends on the offset amount. When the offset amount is 1, the LIR10 unit determines the target data as "01"; and generates control signals for switching node S11 and switching node S10 based on the target data "01". Among them, the LIR10 unit can generate the control signal of the switching node S11 as the first signal based on the "0" in the high bit of "01"; the LIR10 unit can generate the control signal of the switching node S10 as the second signal based on the "1" in the low bit of "01".

[0099] In some embodiments, for any LIR unit, the LIR unit is used to obtain a second element corresponding to the position of the switching node from a second type of elements among multiple elements based on the position of the switching node in the sub-network connected to the LIR unit, and to determine an offset based on the state value of the second element, wherein the second type of element is in an odd position among multiple elements.

[0100] For example, see continue. Figure 4 The second type of elements among the multiple elements are the elements located at positions 1, 3, 5, and 7, namely elements b, d, f, and h, respectively. The LIR10 unit determines its offset based on the state value 001 of element b at position 1. The LIR11 unit determines its offset based on the state value 011 of element f at position 5. The LIR20 unit determines its offset based on the state value 010 of element d at position 3.

[0101] Specifically, for any LIR unit, the LIR unit is used to extract a target number of low-order bits from the state value of the second element corresponding to the LIR unit, using these bits as offsets. The target number is related to the level of the sub-network corresponding to the LIR unit. For example, the LIR10 unit extracts the lower two bits "01" from the state value 001 of element b. Binary "01" equals decimal 1, meaning the offset of the LIR10 unit is 1. The LIR11 unit extracts the lower two bits "11" from the state value 011 of element f. Binary "11" equals decimal 3, meaning the offset of the LIR10 unit is 3.

[0102] The processing circuit for non-first-level sub-networks is constructed based on a tree structure. The processing circuit includes multiple input terminals and multiple output terminals. The multiple input terminals are used to input the prompt information corresponding to each element in multiple vectors. The multiple output terminals are connected to multiple LIR units. For example, Figure 7 This is a schematic diagram of another processing circuit provided according to an embodiment of this application. The state value of each element is calculated using a tree structure of pairwise addition. Due to the relatively large delay of the adder, this tree structure has a smaller delay compared to a chained addition structure. Finally, the target number of values ​​are extracted at the corresponding positions as the output state values ​​of the first, third, and fifth elements. It should be noted that only the lower 2 bits of data are extracted for the state values ​​of the first and fifth elements, while the lower 3 bits are extracted for the state value of the third element.

[0103] This application provides a processor that uses a butterfly network to aggregate valid elements from multiple vectors. Since the butterfly network performs data selection hierarchically, with the number of levels and the length of the input vector having a logarithmic relationship, the layout area of ​​the butterfly network is significantly reduced. Furthermore, each sub-network uses a pairwise aggregation method with swapping nodes to exchange the positions of elements in multiple vectors, resulting in lower wiring complexity. This not only saves overall resource consumption and layout costs but also facilitates faster acquisition of valid elements from the vectors. Moreover, by connecting a switch control circuit to each swapping node in the butterfly network, precise control of each swapping node is achieved, thereby more accurately controlling the flow direction of elements in the vectors and more accurately acquiring valid elements, thus improving processor performance.

[0104] The aforementioned processor can be deployed in any computer device, whether a terminal or a server, enabling that computer device to execute the vector convergence method based on the processor. The following section first uses a computer device as an example to introduce the implementation environment of the vector convergence method provided in this application embodiment. Figure 8 This is a schematic diagram illustrating the implementation environment of a vector convergence method according to an embodiment of this application. See also... Figure 8 The implementation environment includes a terminal 801 and a server 802. The terminal 801 and the server 802 can be connected directly or indirectly via wired or wireless communication, which is not limited herein.

[0105] In some embodiments, terminal 801 is a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, smart voice interaction device, smart home appliance, in-vehicle terminal, etc., but is not limited thereto. Terminal 801 is equipped with the aforementioned processor. Illustratively, terminal 801 is a terminal used by a user. Terminal 801 can obtain multiple vectors from server 802. Then, terminal 801 executes the vector aggregation method provided in this application through the processor to aggregate the valid information (valid elements) from the multiple vectors.

[0106] Those skilled in the art will understand that the number of terminals described above can be more or less. For example, there may be only one terminal, or there may be dozens or hundreds of terminals, or even more. This application does not limit the number of terminals or the type of device.

[0107] In some embodiments, server 802 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), big data, and artificial intelligence platforms. Server 802 can provide multiple vectors to terminal 801. Alternatively, server 802 executes the vector aggregation method provided in the embodiments of this application, which is not limited in this embodiment. In some embodiments, server 802 undertakes the main computing work, and terminal 801 undertakes the secondary computing work; or, server 802 undertakes the secondary computing work, and terminal 801 undertakes the main computing work; or, server 802 and terminal 801 collaborate on computing using a distributed computing architecture.

[0108] Figure 9 This is a flowchart of a vector convergence method provided according to an embodiment of this application. See also... Figure 9 In this embodiment, the vector convergence method is described using an example executed by a processor in a computer device. The vector convergence method includes the following steps:

[0109] 901. In response to the vector convergence instruction, the processor generates control signals for each switching node in the butterfly network within the processor based on prompts from multiple vectors. The prompts are used to indicate the valid elements in the multiple vectors.

[0110] 902. The processor inputs multiple vectors into the butterfly network and controls the flow of multiple elements in the multiple vectors in the butterfly network based on the control signals of each switching node in the butterfly network.

[0111] 903. For any switching node, when the control signal of the switching node is the first signal, the processor performs position swapping on the two elements input in the switching node; when the control signal is the second signal, it does not perform position swapping on the two elements input in the switching node.

[0112] 904. The processor outputs a target vector, which is composed of the valid elements of multiple vectors.

[0113] The principle of the vector convergence method in steps 901 to 904 is the same as the principle of vector convergence in the various parts of the processor described above, and will not be repeated here. That is, the vector convergence method provided in the above embodiments and the processor embodiments belong to the same concept, and its specific implementation process can be found in the processor embodiments, which will not be repeated here.

[0114] Any vector processed by the processor can be a feature vector obtained based on certain data, such as an image feature vector or a text feature vector, or it can be a vector composed of multiple values ​​in any array (equivalent to the array). This application does not limit this.

[0115] Taking an example where multiple input vectors in a butterfly network are all image feature vectors, the vector convergence method provided in this application can be applied to image fusion scenarios. Accordingly, the processor generates control signals for each exchange node in the butterfly network based on the prompt information of multiple image feature vectors. The prompt information for each image feature vector is used to indicate the effective element in that image feature vector. That is, the prompt information for each image feature vector is used to indicate the effective region in the image represented by that image feature vector. The processor inputs multiple elements from multiple image feature vectors into the butterfly network, and controls the flow of multiple elements from multiple image feature vectors in the butterfly network based on the control signals of each exchange node in the butterfly network. For any exchange node, when the control signal of the exchange node is the first signal, the processor swaps the positions of the two elements input to the exchange node; when the control signal is the second signal, it does not swap the positions of the two elements input to the exchange node. Then, the processor outputs the target image feature vector through the butterfly network. The target image feature vector is used to represent the fused image. That is, the prompt information for each image is used to indicate the effective region (foreground region or background region) in the image. The above method can be used to merge the foreground area of ​​one image with the background area of ​​another image to obtain a new image.

[0116] Taking a butterfly network where multiple input vectors are text feature vectors as an example, the vector convergence method provided in this application can be used to extract key information from text. Accordingly, the processor generates control signals for each exchange node in the butterfly network based on the prompt information of multiple text feature vectors. The prompt information for each text feature vector represents a valid element in that text feature vector. That is, the prompt information for each text feature vector represents a keyword (valid information) in the sentence represented by that text feature vector. The processor inputs multiple elements from multiple text feature vectors into the butterfly network and controls the flow of these elements within the butterfly network based on the control signals of each exchange node. For any exchange node, if the control signal of the exchange node is a first signal, the processor swaps the positions of two elements input into the exchange node; if the control signal is a second signal, it does not swap the positions of the two elements input into the exchange node. Then, the processor outputs a target text feature vector through the butterfly network. The target text feature vector represents key information in the text. This key information can be a summary of the text, and this application does not limit this.

[0117] Taking the example of multiple vectors used as input in a butterfly network to represent arrays, each array contains first-class, second-class, and third-class information. The processor can use the vector aggregation method described above to aggregate information of a specific class from multiple arrays. The principle of this process is similar to that of aggregating image or text information, and will not be elaborated further here. For example, each array represents information about a specific user. For any given array, this array includes information such as the user's gender, age, and grades. The processor can use the vector aggregation method described above to aggregate the grades of multiple users to view the learning progress of multiple users, etc.

[0118] This application provides a vector aggregation method that uses a butterfly network to aggregate valid elements from multiple vectors. Since the butterfly network performs data selection hierarchically, with the number of levels and the length of the input vector having a logarithmic relationship, the layout area of ​​the butterfly network is significantly reduced. Furthermore, each sub-network uses a pairwise aggregation method with swapping nodes to exchange the positions of elements in multiple vectors, resulting in lower wiring complexity. This not only saves overall resource consumption and layout costs but also facilitates faster acquisition of valid elements from the vectors. Moreover, by controlling each swapping node separately, the flow direction of elements in the vectors can be controlled more accurately, leading to more accurate acquisition of valid elements and improved processor performance.

[0119] In the embodiments of this application, the computer device can be configured as a terminal or a server. When the computer device is configured as a terminal, the terminal can act as the execution subject to implement the technical solutions provided in the embodiments of this application. When the computer device is configured as a server, the server can act as the execution subject to implement the technical solutions provided in the embodiments of this application. Alternatively, the technical solutions provided in this application can be implemented through the interaction between the terminal and the server. The embodiments of this application do not limit this.

[0120] Figure 10 This is a structural block diagram of a terminal 1000 provided according to an embodiment of this application. The terminal 1000 can be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The terminal 1000 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other names.

[0121] Typically, terminal 1000 includes a processor 1001 and a memory 1002.

[0122] Processor 1001 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1001 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1001 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1001 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0123] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 are used to store at least one computer program, which is executed by the processor 1001 to implement the vector convergence method provided in the method embodiments of this application.

[0124] In some embodiments, the terminal 1000 may also optionally include a peripheral device interface 1003 and at least one peripheral device. The processor 1001, memory 1002, and peripheral device interface 1003 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1003 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1004, a display screen 1005, a camera assembly 1006, an audio circuit 1007, and a power supply 1008.

[0125] Peripheral device interface 1003 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1001 and memory 1002. In some embodiments, processor 1001, memory 1002 and peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1001, memory 1002 and peripheral device interface 1003 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0126] The radio frequency (RF) circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1004 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1004 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. In some embodiments, the RF circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1004 can communicate with other terminals via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1004 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0127] Display screen 1005 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1005 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1001 for processing. In this case, display screen 1005 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1005, disposed on the front panel of terminal 1000; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal 1000 or in a folded design; in still other embodiments, display screen 1005 may be a flexible display screen, disposed on a curved or folded surface of terminal 1000. Furthermore, display screen 1005 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1005 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0128] The camera assembly 1006 is used to acquire images or videos. In some embodiments, the camera assembly 1006 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1006 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash is a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0129] The audio circuit 1007 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1001 for processing, or input to the radio frequency circuit 1004 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 1000. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1001 or the radio frequency circuit 1004 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1007 may also include a headphone jack.

[0130] The power supply 1008 is used to power the various components in the terminal 1000. The power supply 1008 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When the power supply 1008 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0131] In some embodiments, the terminal 1000 further includes one or more sensors 1009. The one or more sensors 1009 include, but are not limited to: an acceleration sensor 1010, a gyroscope sensor 1011, a pressure sensor 1012, an optical sensor 1013, and a proximity sensor 1014.

[0132] Accelerometer 1010 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal 1000. For example, accelerometer 1010 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1001 can control display screen 1005 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1010. Accelerometer 1010 can also be used for games or for acquiring user motion data.

[0133] The gyroscope sensor 1011 can detect the orientation and rotation angle of the terminal 1000. The gyroscope sensor 1011 can work in conjunction with the accelerometer sensor 1010 to collect the user's 3D movements on the terminal 1000. Based on the data collected by the gyroscope sensor 1011, the processor 1001 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0134] The pressure sensor 1012 can be disposed on the side bezel of the terminal 1000 and / or on the lower layer of the display screen 1005. When the pressure sensor 1012 is disposed on the side bezel of the terminal 1000, it can detect the user's grip signal on the terminal 1000, and the processor 1001 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1012. When the pressure sensor 1012 is disposed on the lower layer of the display screen 1005, the processor 1001 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1005. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0135] An optical sensor 1013 is used to collect ambient light intensity. In one embodiment, the processor 1001 can control the display brightness of the display screen 1005 based on the ambient light intensity collected by the optical sensor 1013. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1005 is increased; when the ambient light intensity is low, the display brightness of the display screen 1005 is decreased. In another embodiment, the processor 1001 can also dynamically adjust the shooting parameters of the camera assembly 1006 based on the ambient light intensity collected by the optical sensor 1013.

[0136] The proximity sensor 1014, also known as a distance sensor, is typically installed on the front panel of the terminal 1000. The proximity sensor 1014 is used to detect the distance between the user and the front of the terminal 1000. In one embodiment, when the proximity sensor 1014 detects that the distance between the user and the front of the terminal 1000 is gradually decreasing, the processor 1001 controls the display screen 1005 to switch from a screen-on state to a screen-off state; when the proximity sensor 1014 detects that the distance between the user and the front of the terminal 1000 is gradually increasing, the processor 1001 controls the display screen 1005 to switch from a screen-off state to a screen-on state.

[0137] Those skilled in the art will understand that Figure 10 The structure shown does not constitute a limitation on terminal 1000 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0138] Figure 11This is a schematic diagram of a server structure according to an embodiment of this application. The server 1100 can vary considerably due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 1101 and one or more memories 1102. The memory 1102 stores at least one computer program, which is loaded and executed by the processor 1101 to implement the vector aggregation method provided in the above-described method embodiments. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated here.

[0139] This application also provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor of a computer device to implement the operations performed by the computer device in the vector convergence method of the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, or optical data storage device, etc.

[0140] In some embodiments, the computer program involved in the present application embodiments may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network may constitute a blockchain system.

[0141] This application also provides a computer program product, including a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the vector convergence method provided in the various optional implementations described above.

[0142] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0143] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A processor, characterized in that, The processor includes: a butterfly network and a switch control circuit; The butterfly network includes multiple levels of sub-networks, and each level of sub-network includes multiple switching nodes; In any sub-network, the multiple switching nodes in the sub-network are not connected; For the i-th level subnetwork and the (i+1)-th level subnetwork, the j-th switching node in the i-th level subnetwork is connected to the j-th and n-th switching nodes in the (i+1)-th level subnetwork, respectively. The n-th switching node in the i-th level subnetwork is connected to the n-th and j-th switching nodes in the (i+1)-th level subnetwork, respectively. The difference between n and j is 2. i-1 ; The switch control circuit is connected to each switching node in the butterfly network respectively; The switch control circuit is used to respond to a vector convergence command and generate control signals for each switching node in the butterfly network based on prompt information from multiple vectors, wherein the prompt information is used to indicate the valid elements in the multiple vectors; Inputting multiple elements from the multiple vectors into the first-level subnetwork of the butterfly network; For any switching node in the butterfly network, the switching node is used to swap the positions of two elements input to the switching node when the control signal is a first signal, and not to swap the positions of the two elements input to the switching node when the control signal is a second signal. The last sub-network in the butterfly network outputs a target vector, which is composed of the valid elements of the plurality of vectors.

2. The processor according to claim 1, characterized in that, The switch control circuit includes multiple switch control modules, and each switch control module corresponds one-to-one with the multi-level sub-network. For any switch control module, the switch control module is connected to the switching node in the corresponding sub-network; the switch control module is used to respond to the vector convergence command, determine the state value of the multiple elements based on the prompt information of the multiple vectors, and generate the control signal of the corresponding switching node in the sub-network based on the state value of the multiple elements. The state value of each element is used to represent the number of valid elements between the element and the target element, and the target element is the least significant element among the multiple elements.

3. The processor according to claim 2, characterized in that, The switch control module corresponding to the first-level sub-network in the butterfly network includes multiple LSB units, and each of the multiple LSB units corresponds to a switching node in the first-level sub-network. For any LSB unit, the LSB unit is connected to the control terminal of the corresponding switching node, and the control terminal is used to receive the control signal generated by the LSB unit; The LSB unit is configured to, based on the position of the switching node to which the LSB unit is connected in the primary sub-network, obtain a first element corresponding to the position of the switching node from a first type of elements among the plurality of elements, generate a first signal when the least significant bit of the state value of the first element is a first value, and generate a second signal when the least significant bit of the state value of the first element is a second value, wherein the first type of element is in an even position among the plurality of elements, and the state value of each element among the plurality of elements is a binary value.

4. The processor according to claim 3, characterized in that, The switch control module corresponding to the primary sub-network in the butterfly network also includes a processing circuit. The processing circuit is constructed based on the form of an XOR chain. The processing circuit includes multiple input terminals and multiple output terminals. The multiple input terminals are used to input the prompt information corresponding to each element in the multiple vectors respectively. The multiple output terminals are connected to the multiple LSB units respectively. The multiple output terminals are used to provide the least significant bit of the state value of the first type of element to the multiple LSB units.

5. The processor according to claim 2, characterized in that, The switch control module corresponding to the non-primary subnet in the butterfly network includes at least one LIR unit. For any LIR unit, the LIR unit is connected to the control terminal of the corresponding switching node, and the control terminal is used to receive the control signal generated by the LIR unit; The LIR unit is used to perform an offset operation on preset data based on the offset of the LIR unit to obtain target data, and generate a control signal based on the target data. The preset data is related to the level of the sub-network corresponding to the LIR unit, and the offset is the number of times the offset operation is performed. In any offset operation, the LIR unit is used to extract the value of the highest bit in the preset data, shift the remaining value in the preset data one bit to the left, invert the value of the highest bit, and place the inverted value in the lowest bit of the preset data.

6. The processor according to claim 5, characterized in that, For any LIR unit, the LIR unit is used to obtain a second element corresponding to the position of the switching node from a second type of elements among the plurality of elements based on the position of the switching node in the sub-network connected to the LIR unit, and to determine the offset based on the state value of the second element, wherein the second type of element is in an odd position among the plurality of elements.

7. The processor according to claim 6, characterized in that, For any LIR unit, the LIR unit is used to obtain a target number of low-order values ​​from the state value of the second element corresponding to the LIR unit, and use the values ​​as the offset. The target number is related to the level of the sub-network corresponding to the LIR unit.

8. A vector convergence method, characterized in that, The method is executed by a processor, wherein the processor is the processor described in any one of claims 1 to 7; The method includes: In response to a vector convergence command, control signals for each switching node in the butterfly network of the processor are generated based on prompt information from multiple vectors, wherein the prompt information is used to indicate the valid elements in the multiple vectors; The plurality of vectors are input into the butterfly network, and based on the control signals of each switching node in the butterfly network, the flow of multiple elements in the plurality of vectors in the butterfly network is controlled. For any switching node, if the control signal of the switching node is the first signal, the two elements input in the switching node are swapped; if the control signal is the second signal, the two elements input in the switching node are not swapped. Output a target vector, which is composed of the valid elements of the plurality of vectors.

9. A computer device, characterized in that, The computer device includes a processor and a memory, the processor being the processor according to any one of claims 1 to 7, and the memory being used to store at least one computer program, the at least one computer program being loaded by the processor and executed as the vector convergence method according to claim 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one computer program for executing the vector convergence method of claim 8.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the vector convergence method as described in claim 8, wherein the processor is the processor described in any one of claims 1 to 7.