Information processing method and device, equipment and storage medium
By segmenting the input features into multiple feature components and combining dynamic weight processing, the gradient vanishing and memory overhead problems in residual connections are solved, and the efficiency and performance of the model are improved.
Patent Information
- Application Number
- CN202510323320.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-01
AI Technical Summary
Existing residual connections have gradient vanishing and representation crash problems in deep learning, and the introduction of multiple connection methods leads to excessive memory overhead.
The first input feature is divided into multiple feature components, and these components are processed through weight parameters to generate intermediate input features and output features, and in place of feature copying and expanding width, combined with dynamic adjustment of connection weights, solve the problems of gradient vanishing and memory overhead.
It reduces memory usage and computing costs, while improving the processing efficiency and generalization capabilities of the model, and improving the performance of the network model.
Smart Images

Figure CN120234529A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for information processing. Background Art
[0002] With the rapid development of deep learning technology, the emergence of the Residual Connection technology is of great significance. The basic idea of the Residual Connection is to add shortcut connections between some layers of the neural network, so that the network can directly learn the residual mapping between the input and output, rather than directly learning the target function. This structure helps to alleviate the vanishing gradient problem and makes the training of deep networks easier and more effective. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for information processing is provided. The method includes: providing input information to a target model, where the target model includes a plurality of processing layers, and the plurality of processing layers at least include adjacent first and second processing layers; splitting a first input feature of the first processing layer into a first set of feature components, where the first input feature is determined based on the input information; applying a first set of weight parameters to the first set of feature components to determine an intermediate input feature; determining an intermediate output feature generated by the first processing layer based on the intermediate input feature; determining a second set of feature components based on the intermediate output feature and the first set of feature components as a second input feature for the second processing layer; and generating an output result of the target model based at least on the second input feature.
[0004] In a second aspect of the present disclosure, an apparatus for information processing is provided. The apparatus includes: an input providing module configured to provide input information to a target model, where the target model includes a plurality of processing layers, and the plurality of processing layers at least include adjacent first and second processing layers; a feature splitting module configured to split a first input feature of the first processing layer into a first set of feature components, where the first input feature is determined based on the input information; a first weighting module configured to apply a first set of weight parameters to the first set of feature components to determine an intermediate input feature; a feature processing module configured to determine an intermediate output feature generated by the first processing layer based on the intermediate input feature; a second weighting module configured to determine a second set of feature components based on the intermediate output feature and the first set of feature components as a second input feature for the second processing layer; and a result output module configured to generate an output result of the target model based at least on the second input feature.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium, and the computer program is executable by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of the present disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method according to the first aspect of the present disclosure.
[0008] It should be understood that the content described in this part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0010] Figure 1 A schematic diagram showing a conventional residual connection is shown;
[0011] Figure 2 A schematic diagram showing an example process of information processing according to some embodiments of the present disclosure is shown;
[0012] Figure 3 A schematic diagram showing an example connection architecture according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A schematic structural block diagram showing an example device for information processing according to some embodiments of the present disclosure is shown;
[0014] Figure 5 A block diagram of an electronic device capable of implementing multiple embodiments of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0016] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. In addition, the embodiments described in any section / subsection can be combined with any other embodiments described in the same section / subsection and / or different section / subsections in any manner.
[0017] In the description of the embodiments of the present disclosure, the term "comprising" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter. The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0018] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects all comply with the corresponding laws, regulations and related provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware and confirms. Accordingly, when implementing the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the data or information that may be involved should be informed to the user and the user's authorization should be obtained through appropriate means according to the relevant laws and regulations. The specific informing and / or authorization methods may vary according to the actual situation and application scenarios, and the scope of the present disclosure is not limited in this regard.
[0019] In the solutions of this specification and embodiments, if personal information processing is involved, it will be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for performing a contract, etc.), and will only be processed within the specified or agreed scope. If the user refuses to process personal information other than the necessary information required for basic functions, it will not affect the user's use of basic functions.
[0020] The residual connection enables the network to more effectively learn the residual mapping between the input data and the output data by adding a skip connection to each layer of the neural network, rather than directly learning the mapping itself.Figure 1 FIG. 1 shows a schematic diagram of a traditional residual connection. As shown in the figure, after processing the input feature X by the processing layer to perform F(X), the input feature X can be added to the output feature F(X) as the input feature of the next processing layer. However, the residual connection is affected by the vanishing gradient and representation collapse. Therefore, the prior art proposes to introduce multiple connections at different network depths to solve the problems of vanishing gradient and representation collapse. However, introducing multiple connections at different network depths leads to excessive memory overhead.
[0021] Embodiments of the present disclosure propose a scheme for information processing. The scheme includes: providing input information to a target model, where the target model includes multiple processing layers, and the multiple processing layers at least include adjacent first and second processing layers; splitting the first input feature of the first processing layer into a first set of feature components, where the first input feature is determined based on the input information; applying the first set of weight parameters to the first set of feature components to determine an intermediate input feature; determining the intermediate output feature generated by the first processing layer based on the intermediate input feature; determining a second set of feature components based on the intermediate output feature and the first set of feature components as the second input feature for the second processing layer; and generating an output result of the target model based at least on the second input feature.
[0022] In this way, embodiments of the present disclosure can, by splitting the first input feature into a first set of feature components, replace the method of copying the first input feature multiple times, introducing learnable connection weights in the neural network. On the premise of ensuring that a more powerful and flexible connection mechanism is provided for the model, helping to solve the representation collapse caused by the residual connection, and improving the efficiency of model processing, the memory overhead is reduced.
[0023] Various example implementations of the scheme are further described in detail below with reference to the accompanying drawings.
[0024] Figure 2 FIG. 2 shows a flowchart of an example information processing process 200 according to some embodiments of the present disclosure. Process 200 can be implemented at a suitable electronic device, and such an electronic device can be deployed with a target model.
[0025] As Figure 2 shown, at block 210, the electronic device provides input information to the target model, where the target model includes multiple processing layers, and the multiple processing layers at least include adjacent first and second processing layers.
[0026] Figure 3 FIG. 3 shows an example connection architecture 300 according to some embodiments of the present disclosure. As Figure 3As shown, the architecture of the target model may include multiple processing layers. For example, the attention layer 340 and the feed-forward layer 370. In some embodiments, the target model may be constructed by replacing a set of processing layers connected by residual connections in the transformer unit with multiple processing layers connected by the connection methods introduced below.
[0027] In some embodiments, the target model may also be a generative model, and the input information received by the target model may include text content. The text content may be input as a prompt for the generative model.
[0028] Such a generative model can be used to perform appropriate types of generation tasks. For example, generate text content, image content (pictures or videos), audio content, etc. based on the input prompt.
[0029] At block 220, the electronic device splits the first input feature of the first processing layer into a first set of feature components, and the first input feature is determined based on the input information.
[0030] For Figure 3 example, the electronic device 110 may use the target model to process the input information. Figure 3 In this case, the attention layer 340 is the first processing layer, and may determine the first input feature 310 of the first processing layer, and split the first input feature 310 into a first set of feature components. As Figure 3 shown, the first set of feature components includes the feature component 320 and the feature component 330.
[0031] In some embodiments, the electronic device 110 may split the first input feature into a first set of feature components based on a preset number of components or a preset feature length. As an example, h k-1 can be used to represent the input feature associated with the k-th layer. The initial input h 0 of the network may be split into m parts to form an initial hidden feature matrix Reshape means changing the dimensional structure of the input feature without changing its data content, making the input feature more convenient for subsequent processing steps of the network model. Accordingly, the input feature matrix of the k-th processing layer can be expressed as where represents the i-th feature component in the input feature of the k-th processing layer.
[0032] By splitting the input feature into multiple feature components instead of expanding the feature width through feature replication, the embodiments of the present disclosure can avoid the additional memory overhead caused by the increase in feature width, thereby reducing memory usage and lowering the computational cost.
[0033] At block 230, the electronic device applies a first set of weight parameters to the first set of feature components to determine intermediate input features.
[0034] In some embodiments, the electronic device 110 may determine second input features for the second processing layer by concatenating a second set of feature components. As Figure 3 shown, taking the first input feature 310 including the feature component 320 and the feature component 330 as an example, the first input feature 310 of the first processing layer may be determined based on the weighted sum of the feature component 320 and the feature component 330. As an example, the intermediate input features may include two parts: and For and perform feature concatenation processing to obtain intermediate input features where and are the first set of weight parameters. The number of parameters of the first set of weight parameters is associated with the number of components of the first set of feature components. For example, if the first input feature is divided into m feature components, then the number of parameters of the first set of weight parameters is m×m.
[0035] As an example, the input feature of the k-th processing layer can be expressed as and the intermediate input feature of the k-th processing layer can be expressed as: where Y represents the first set of weight parameters
[0036] Continuing to refer to Figure 2 , at block 240, the electronic device determines the intermediate output features generated by the first processing layer based on the intermediate input features.
[0037] As an example, the first processing layer may perform processing on the intermediate input features (for example, this process can be expressed as T), and the corresponding intermediate output features can be expressed as
[0038] At block 250, the electronic device determines a second set of feature components based on the intermediate output features and the first set of feature components as the second input features for the second processing layer.
[0039] Continuing to refer to Figure 3 , the electronic device may determine the second set of feature components of the second input features based on the intermediate output features and the first set of feature components The second set of feature components includes the feature component 350 and the feature component 360, where each feature component of the second set of feature components is determined by applying the corresponding second set of weight parameters to the intermediate output features and the first set of feature components.
[0040] Taking Figure 3 as an example, the first feature component of the second input feature can be calculated based on and The second feature component of the second input feature can be calculated based on and .
[0041] Thus, the embodiments of the present disclosure consider both the width connection of different feature components of the first input feature and the depth connection between the first input feature and the intermediate output feature of the processing layer.
[0042] The determination process of the second input feature can also be represented as a matrix operation, which can process the first input feature based on the matrix FC of the weight parameters to determine the second input feature. As an example, FC can be represented as:
[0043]
[0044] The processing process of the second input feature can be represented as:
[0045]
[0046] In some embodiments, the weight parameters included in the matrix FC of the weight parameters can be static parameters determined by training the target model. Taking Figure 3 as an example, the first set of weight parameters and / or the second set of weight parameters, for example can be static model parameters determined by training the target model.
[0047] In some embodiments, the number of parameters of the static parameter matrix FC can be determined according to the following formula:
[0048] |θ SHC | = |θ B | + |θ Y | + |θ A | = m + m×m + m×m = m×(2m + 1) (3)
[0049] The number of additional parameters is:
[0050] P extra = |θ SHC |×2×(k - 1) (4)
[0051] In some embodiments, the first set of weight parameters and / or the second set of weight parameters can also be dynamic parameters determined according to the first input feature.
[0052] As an example, the matrix of weight parameters can be expressed as:
[0053]
[0054] Correspondingly, the process of determining the second input feature based on the first input feature can be expressed as:
[0055]
[0056] In some embodiments, the electronic device can dynamically determine the weight parameters based on the following process:
[0057]
[0058] Specifically, as shown in formula (7), the electronic device can normalize the first input feature H to determine the reference feature
[0059] Furthermore, as represented by formulas (8) to (10), the electronic device can determine the training parameters determined by training the target model. Specifically, the electronic device can determine the training parameters s β , s α , s γ , W β , W α , W γ , where B, A, and Y are preset parameters.
[0060] Furthermore, the electronic device can perform a linear transformation on the reference feature according to formulas (8) to (10). Specifically, the electronic device can perform a linear transformation on the reference feature based on the training parameters to determine the first set of weight parameters and / or the second set of weight parameters, thereby determining the dynamic weight parameter matrix FC(H).
[0061] In some embodiments, the number of parameters of the dynamic parameter matrix can be determined according to the following formula:
[0062]
[0063] d model is the dimension of the hidden state in the network model, and |θ norm | depends on the type of the normalization module. Then the number of additional parameters is:
[0064] P extra = |θ DFC | × 2 × (k - 1) (12)
[0065] In addition, the computational complexity of applying the weight parameter FC in the fully connected layer to implement feature transfer is O(dmodel × 2m), the computational complexity of the feedforward layer is O(2 × d model × d ffn ), and the computational complexity of the attention layer is O(4 × d model × d model ). Since O(d model × 2m) << O(4 × d model × d model ) << O(2 × d model × d ffn ), the computational cost required in the fully connected layer can be ignored compared with the feedforward layer and the attention layer. Therefore, the computational overhead introduced by applying the weight parameters to achieve feature transfer is extremely small and will not affect the overall computational cost.
[0066] The dynamic weight parameters can dynamically adjust the weights of different paths according to the input features, enabling the network to more flexibly fuse features, thereby enhancing the adaptability of the network model to different input features. The dynamic weight parameters also help to alleviate the vanishing gradient problem to ensure the effective propagation of gradients. Moreover, since the computational overhead increased by the dynamic weight parameters is almost negligible, the utilization efficiency of computational resources can be improved. Additionally, since the dynamic weight parameters can be adjusted by themselves according to the task requirements, the network model performs better in different tasks, enhancing the generalization ability of the network model. Based on the above methods, the embodiments of the present disclosure can bring significant performance improvement and computational efficiency to the network model, while reducing the burden on the designer in network architecture design.
[0067] Continuing to refer to Figure 2 , at block 260, the electronic device generates the output result of the target model based at least on the second input feature.
[0068] As an example, the second processing layer can process the second input feature based on the architecture described in Figure 3 and determine the input feature of the next processing layer considering the depth connection and the width connection.
[0069] In some embodiments, an initialization strategy can also be performed on the embodiments of the present disclosure. For example, the training parameters W β , W α , W γ can be initialized to 0, and the static parameters can be initialized as follows:
[0070]
[0071] In some embodiments, referring to Figure 3 , the processing layer can be the feedforward layer 370 or the attention layer 340, Figure 3The middle attention layer 340 can be the first processing layer, and the feed-forward layer 370 can be the second processing layer. The electronic device can provide input information to the target model. Then, based on the input information, a first input feature 310 associated with the attention layer 340 is determined. The first input feature 310 includes a feature component 320 and a feature component 330. The first input feature 310 can be expressed as where represents the i-th feature component in the first input feature 310 of the attention layer 340.
[0072] The first input feature 310 of the attention layer 340 can be determined according to the weighted sum of the feature component 320 and the feature component 330. The intermediate input feature can include two parts: and Then, perform feature splicing processing on and to obtain the intermediate input feature where and are the first set of weight parameters. Next, apply the attention layer 340 to process the intermediate input feature to obtain an intermediate output feature
[0073] Then, based on the intermediate output feature and the first set of feature components determine the second set of feature components of the second input feature. The second set of feature components includes a feature component 350 and a feature component 360. The first feature component of the second input feature can be calculated based on and The second feature component of the second input feature can be calculated based on and
[0074] Then, determine the second input feature associated with the feed-forward layer 370. The second input feature includes a feature component 350 and a feature component 360. The second input feature can be expressed as where represents the i-th feature component in the second input feature of the feed-forward layer 370.
[0075] The second input feature of the feed-forward layer 370 can be determined according to the weighted sum of the feature component 350 and the feature component 360. The intermediate input feature can include two parts: and Then, for and perform feature splicing processing to obtain intermediate input features where and are the first set of weight parameters. Next, apply the feedforward layer 370 to process the intermediate input features to obtain intermediate output features
[0076] Then, based on the intermediate output features and the second set of feature components determine the third set of feature components of the third input feature. The first feature component of the third input feature can be calculated based on and The second feature component of the third input feature can include: and Then, according to the architecture described in the reference Figure 3 process the third input feature and determine the input features of the third processing layer.
[0077] Furthermore, the target model can determine the final output result of the model through the processing layer connection architecture described above. As mentioned above, the target model can be a generative model. Correspondingly, its input results can include but are not limited to: text content, image content (pictures or videos), audio content, etc.
[0078] By splitting the first input feature into the first set of feature components instead of expanding the feature width through feature replication, the embodiments of the present disclosure can avoid the additional memory overhead caused by the increase in feature width, thereby reducing memory usage and computational cost. In addition, by dynamically adjusting the connection weights between different layers, the dilemma of the seesaw trade-off between gradient disappearance and representation collapse is solved. Moreover, the embodiments of the present disclosure can also significantly improve the performance of the model, including language modeling, image classification, and diffusion models, etc., demonstrating its wide applicability and effectiveness.
[0079] Experiments show that for the scenario where the input features are not split, as the number of tokens increases, the training loss will decrease. However, for the scenario where the input features are split into two feature vectors, as the number of tokens increases, the degree of decrease in the training loss is more obvious. This shows that the method of splitting the input features and then processing them proposed by the present disclosure can significantly optimize the output quality of the network model.
[0080] Example devices and equipment
[0081] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 4FIG. 0 shows a schematic structural block diagram of an example apparatus 400 for information processing according to certain embodiments of the present disclosure. The apparatus 400 may be implemented as or included in an electronic device as discussed above. Each module / component in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0082] As Figure 4 shown, the apparatus 400 includes an input providing module 410 configured to provide input information to a target model, the target model including a plurality of processing layers, the plurality of processing layers including at least adjacent first and second processing layers; a feature segmentation module 420 configured to segment a first input feature of the first processing layer into a first set of feature components, the first input feature being determined based on the input information; a first weighting module 430 configured to apply a first set of weight parameters to the first set of feature components to determine an intermediate input feature; a feature processing module 440 configured to determine an intermediate output feature generated by the first processing layer based on the intermediate input feature; a second weighting module 450 configured to determine a second set of feature components based on the intermediate output feature and the first set of feature components as a second input feature for the second processing layer; and a result output module 460 configured to generate an output result of the target model based at least on the second input feature.
[0083] In some embodiments, the feature segmentation module 420 is further configured to segment the first input feature into the first set of feature components based on a preset number of components or a preset feature length.
[0084] In some embodiments, the second weighting module 450 is further configured to determine the second input feature for the second processing layer by concatenating the second set of feature components.
[0085] In some embodiments, each feature component of the second set of feature components is determined by applying a corresponding second set of weight parameters to the intermediate output feature and the first set of feature components.
[0086] In some embodiments, the first set of weight parameters and / or the second set of weight parameters are static parameters determined by training the target model.
[0087] In some embodiments, the first set of weight parameters and / or the second set of weight parameters are dynamic parameters determined by the first input feature.
[0088] In some embodiments, the first set of weight parameters and / or the second set of weight parameters are determined based on the following process: normalizing the first input feature to determine a reference feature; determining training parameters determined by training the target model; and performing a linear transformation on the reference feature based on the training parameters to determine the first set of weight parameters and / or the second set of weight parameters.
[0089] In some embodiments, the target model is constructed by replacing a set of processing layers connected by residual connections in the transformer unit with multiple processing layers.
[0090] In some embodiments, the target model is a generative model, and the input information includes at least text content, which is provided as a prompt for the generative model, and the output result includes at least one of the following: text content, image content, audio content.
[0091] In some embodiments, the processing layer is a feed-forward layer or an attention layer.
[0092] The modules included in apparatus 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the modules in apparatus 400 can be at least partially implemented by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0093] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure can be implemented is shown. It should be understood that Figure 5 The illustrated electronic device 500 is merely exemplary and should not impose any limitation on the functions and scope of the embodiments described herein. Figure 5 The illustrated electronic device 500 can be used to implement the electronic device as discussed above.
[0094] As Figure 5 shown, the electronic device 500 is in the form of a general-purpose electronic device. The components of the electronic device 500 can include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 550. The processing unit 510 can be an actual or virtual processor and is capable of performing various processes according to programs stored in the memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 500.
[0095] An electronic device 500 generally includes multiple computer storage media. Such media can be any accessible media that the electronic device 500 can access, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be removable or non-removable media and can include machine-readable media, such as a flash drive, a magnetic disk, or any other media that can be capable of storing information and / or data and can be accessed within the electronic device 500.
[0096] The electronic device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The memory 520 can include a computer program product 525 having one or more program modules that are configured to execute various methods or actions of various embodiments of the present disclosure.
[0097] The communication unit 540 enables communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented by a single computing cluster or multiple computer machines that can communicate through a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.
[0098] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 550 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) as needed through the communication unit 540, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 500, or communicate with any device that enables the electronic device 500 to communicate with one or more other electronic devices (such as a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0099] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is further provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0100] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0101] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0102] The computer-readable program instructions can be loaded onto a computer, other programmable data processing device, or other device, such that a series of operation steps are executed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing device, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0104] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art to understand the various implementations disclosed herein.
Claims
1. A method of information processing, comprising: Providing input information to a target model, the target model comprising a plurality of processing layers, the plurality of processing layers comprising at least a first processing layer and a second processing layer adjacent to each other; Splitting a first input feature of the first processing layer into a first set of feature components, the first input feature being determined based on the input information; Applying a first set of weight parameters to the first set of feature components to determine intermediate input features; determining an intermediate output feature generated by the first processing layer based on the intermediate input feature; determining a second set of feature components based on the intermediate output features and the first set of feature components to serve as second input features for the second processing layer; as well as An output result of the target model is generated based at least on the second input feature.
2. The method of claim 1 , wherein splitting the first input features of the first processing layer into a first set of feature components comprises: The first input feature is segmented into the first group of feature components based on a preset number of components or a preset feature length.
3. The method according to claim 1, further comprising: The second input features for the second processing layer are determined by concatenating the second set of feature components. 4 . The method of claim 1 , wherein each feature component of the second set of feature components is determined by applying a corresponding second set of weight parameters to the intermediate output features and the first set of feature components.
5. The method according to claim 4, wherein the first set of weight parameters and / or the second set of weight parameters are static parameters determined by training the target model. The method according to claim 4 , wherein the first set of weight parameters and / or the second set of weight parameters are dynamic parameters determined based on the first input features.
7. The method according to claim 6, wherein the first set of weight parameters and / or the second set of weight parameters are determined based on the following process: Normalizing the first input feature to determine a reference feature; Determining training parameters determined by training the target model; as well as Based on the training parameters, a linear transformation is performed on the reference features to determine the first set of weight parameters and / or the second set of weight parameters.
8. The method according to claim 1, wherein the target model is constructed by replacing a set of processing layers in a transformer unit via residual connections with the plurality of processing layers.
9. The method according to claim 1, wherein the target model is a generative model, and the input information includes at least text content, the text content is provided as a prompt word for the generative model, and the output result includes at least one of the following: text content, image content, and audio content.
10. The method of claim 1, wherein the processing layer is a feed-forward layer or an attention layer.
11. An apparatus for information processing, comprising: An input providing module configured to provide input information to a target model, wherein the target model includes a plurality of processing layers, wherein the plurality of processing layers includes at least a first processing layer and a second processing layer that are adjacent to each other; a feature segmentation module configured to segment a first input feature of the first processing layer into a first set of feature components, wherein the first input feature is determined based on the input information; a first weighting module configured to apply a first set of weight parameters to the first set of feature components to determine intermediate input features; a feature processing module configured to determine an intermediate output feature generated by the first processing layer based on the intermediate input feature; a second weighting module configured to determine a second set of feature components based on the intermediate output features and the first set of feature components to serve as second input features for the second processing layer; as well as The result output module is configured to generate an output result of the target model based on at least the second input feature.
12. An electronic device comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.
13. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 10.
14. A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.