Information processing method and device, equipment and storage medium

By constructing and compressing feature sequences, combining global and in-group correlation features, and applying attention mechanisms, the problem of high computational complexity and low efficiency of the content recommendation model when processing longer feature sequences is solved, achieving more efficient information processing.

CN120541207APending Publication Date: 2025-08-26DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510599879.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing content recommendation model has high computational complexity, large memory usage and information attenuation when processing long feature sequences, resulting in low information processing efficiency.

Method used

By constructing feature sequences based on target objects, candidate objects and multiple groups of reference objects, using global feature sequences and in-group correlation features, compressing the length of feature sequences, and applying cross attention and self-attention mechanisms to determine the degree of correlation between target objects and candidate objects.

Benefits of technology

It reduces the computational complexity, improves the processing efficiency and expression ability of the model, maintains local semantics and stabilizes the attention distribution, and achieves more efficient information processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541207A_ABST
    Figure CN120541207A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an information processing method and device, equipment and a storage medium. The method comprises the following steps: constructing a first feature sequence based on a target object, a candidate object and multiple groups of reference objects associated with the target object; for each group of reference objects in the plurality of groups of reference objects, updating feature parts corresponding to each group of reference objects in the first feature sequence based on correlation among the reference objects in each group of reference objects so as to determine a second feature sequence; compressing a plurality of feature parts corresponding to the plurality of groups of reference objects in the second feature sequence to determine a third feature sequence, the length of which is smaller than that of the second feature sequence; and based on the third feature sequence, determining an association degree between the target object and the candidate object. In this way, the information processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to information processing methods, apparatuses, devices, and computer-readable storage media. Background Art

[0002] With the development of computer technology, content recommendation models can apply recommendation algorithms to recommend content of interest to users. However, when processing long feature sequences, content recommendation models often face problems such as high computational complexity, large memory usage, and information decay. This results in low information processing efficiency of content recommendation models. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for information processing is provided. The method includes: constructing a first feature sequence based on a target object, candidate objects, and multiple groups of reference objects associated with the target object; updating, for each group of reference objects in the multiple groups of reference objects, a feature portion corresponding to each group of reference objects in the first feature sequence based on correlations between the reference objects in each group of reference objects to determine a second feature sequence; compressing multiple feature portions corresponding to the multiple groups of reference objects in the second feature sequence to determine a third feature sequence, wherein the length of the third feature sequence is less than the length of the second feature sequence; and determining a degree of association between the target object and the candidate objects based on the third feature sequence.

[0004] In a second aspect of the present disclosure, an information processing apparatus is provided. The apparatus includes: a construction module configured to construct a first feature sequence based on a target object, candidate objects, and multiple groups of reference objects associated with the target object; an update module configured to update, for each group of reference objects in the multiple groups of reference objects, a feature portion corresponding to each group of reference objects in the first feature sequence based on correlations between the reference objects in each group of reference objects, to determine a second feature sequence; a compression module configured to compress multiple feature portions corresponding to the multiple groups of reference objects in the second feature sequence to determine a third feature sequence, wherein the length of the third feature sequence is less than the length of the second feature sequence; and a determination module configured to determine a degree of association between the target object and the candidate objects based on the third feature sequence.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided, which includes computer-executable instructions, which, when executed by a processor, implement the method according to the first aspect of the present disclosure.

[0008] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram illustrating an example environment in which embodiments of the present disclosure can be implemented;

[0011] Figure 2 A first flow chart illustrating a process of information processing according to some embodiments of the present disclosure;

[0012] Figure 3 A second flow chart illustrating a process of information processing according to some embodiments of the present disclosure;

[0013] Figure 4 A flow chart of model training according to some embodiments of the present disclosure is shown;

[0014] Figure 5 shows a schematic diagram of a key-value cache according to some embodiments of the present disclosure;

[0015] Figure 6 A schematic structural block diagram of an information processing apparatus according to certain embodiments of the present disclosure is shown;

[0016] Figure 7 A block diagram of an electronic device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0017] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0018] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.

[0019] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.

[0020] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects shall comply with the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms them. Accordingly, when implementing the various embodiments of the present disclosure, the types, scope of use, and usage scenarios of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method may vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.

[0021] If this specification and the solutions in the examples involve the processing of personal information, such processing will be done only with a legitimate basis (such as with the consent of the subject of personal information or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect the user's use of basic functions.

[0022] In content recommendation scenarios, modeling long feature sequences is critical, but existing solutions for processing long feature sequences usually rely on two-stage retrieval or indirect modeling paradigms. Two-stage retrieval refers to first screening out a portion of candidate results from a long feature sequence, and then performing a more detailed evaluation of these candidate results to determine the final recommendation results. Indirect modeling refers to not directly modeling the long feature sequence as a whole, but processing it through some intermediate steps or transformations. Two-stage retrieval or indirect modeling may cause the features or intermediate results extracted upstream to not match the recommended content downstream, and may increase computational complexity, resulting in low efficiency.

[0023] The embodiments of the present disclosure propose an information processing scheme. According to the scheme, a first feature sequence can be constructed based on a target object, candidate objects, and multiple groups of reference objects associated with the target object. Furthermore, for each group of reference objects in the multiple groups of reference objects, based on the correlation between the reference objects in each group of reference objects, the feature parts corresponding to each group of reference objects in the first feature sequence can be updated to determine a second feature sequence. Multiple feature parts corresponding to the multiple groups of reference objects in the second feature sequence can be compressed to determine a third feature sequence, wherein the length of the third feature sequence is less than the length of the second feature sequence. Accordingly, the degree of association between the target object and the candidate object can be determined based on the third feature sequence.

[0024] Based on this approach, the first feature sequence of the embodiment of the present disclosure includes features constructed based on the target object, candidate objects, and multiple groups of reference objects associated with the target object. The construction of this feature sequence can promote the fusion of the overall information of the sequence and stabilize the attention distribution. Feature fusion is then applied to balance efficiency and expressiveness, while retaining local semantics, to achieve global representation by modeling shorter sequences. Next, the feature parts of each group of reference objects are updated based on the correlation between the reference objects, which can maintain the intra-group dependency and reduce the computational cost. By compressing the feature sequence, the computational complexity can be reduced, thereby improving the overall efficiency of the model.

[0025] Sample Environment

[0026] Figure 1 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1As shown, the example environment 100 may include an electronic device 110. In some embodiments, the electronic device 110 may construct a first feature sequence based on a target object, candidate objects, and multiple groups of reference objects associated with the target object; for each group of reference objects in the multiple groups of reference objects, based on the correlation between the reference objects in each group of reference objects, update the feature portion corresponding to each group of reference objects in the first feature sequence to determine a second feature sequence; compress multiple feature portions corresponding to the multiple groups of reference objects in the second feature sequence to determine a third feature sequence, wherein the length of the third feature sequence is less than the length of the second feature sequence; and determine the degree of association between the target object and the candidate object based on the third feature sequence. The above-mentioned information processing method can be implemented by applying the model 120, and the model 120 can be deployed on the electronic device 110, and can also be deployed on other devices, which will not be described in detail here.

[0027] In some embodiments, the electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (Personal Communication System, PCS) device, a personal navigation device, a personal digital assistant (Personal Digital Assistant, PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface for the target user (such as a "wearable" circuit, etc.).

[0028] The electronic device 110 may also be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. The electronic device 110 may include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like.

[0029] It should be understood that the structures and functions of the various embedded elements in the environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.

[0030] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0031] Example Process

[0032] Figure 2 FIG2 is a flowchart showing a process 200 of information processing according to some embodiments of the present disclosure. The process 200 may be implemented at the electronic device 110. Figure 1 Process 200 is described.

[0033] In block 210 , the electronic device 110 constructs a first feature sequence based on the target object, candidate objects, and multiple groups of reference objects associated with the target object.

[0034] In some embodiments, the target object can be a user, or other appropriate physical or virtual object, such as a group, organization, virtual character, digital persona, etc. Candidate objects are objects selected as potential recommended content at the initial stage of the recommendation process. For example, candidate objects can be items, text, media content, etc. In some embodiments, the reference object can be determined based on the target object's historical interaction information. For example, the reference object may include objects with which the target object has historically interacted within a preset period of time.

[0035] When the target object is a user, the acquisition and use of the above historical interaction information is performed with the user's knowledge and authorization. For example, such reference objects may include media content that the user has browsed within a preset time period, and the candidate objects to be recommended may include candidate media content to be recommended to the user.

[0036] When processing longer feature sequences, the model needs to extract key information from the entire feature sequence and associate this information with specific positions or semantic units in the feature sequence to more accurately understand and analyze the feature sequence. Therefore, global features can be added to the feature sequence as auxiliary representations to facilitate the extraction and anchoring of global information. Figure 3 The first feature sequence includes global features 301 and user long sequence features 302. The global features represent the connected context information and sequence features. The global features 301 include two parts, one part 303 is user identification features, context features and cross features, and the other part 304 is candidate object features.

[0037] For example, when processing a piece of text, some contextual information about the text background, text theme, etc. is extracted. This contextual information is then combined with the tags corresponding to each word in the text to form a new input for the model to perform analysis and understanding operations. In this way, the global feature can have a complete attention receptive field, and the model can aggregate the contextual information of the entire feature sequence. At the same time, the global feature acts as a centralized information anchor, enabling information interaction between features in the feature sequence.

[0038] Furthermore, when using the attention mechanism for longer feature sequences, the length of the feature sequence may cause the attention weights to become too dispersed or degraded, or an attention convergence effect may occur, resulting in a loss of focus on key information. However, since global features are special identifiers shared by all sequence positions, their introduction can alleviate the attention convergence effect. This stabilizes the dynamic changes in attention over longer feature sequences even when attention is dispersed, maintaining the diversity of the attention mechanism and ensuring the effectiveness of modeling long-range dependencies.

[0039] In some embodiments, the object set can be represented as U, the target object can be represented as u∈U, the candidate object set can be represented as I, and the sequence generated by the interaction between the target object and the reference object can be represented as The reference object comes from the candidate object set, i.e. The basic features ud corresponding to the target object may include user identification features, context features, cross features, and target items v∈I from the candidate object set, etc.

[0040] In some embodiments, L can be used to represent the sequence length and d can be used to represent the embedding dimension in the feature sequence. When using ordinary transformers to process longer feature sequences, the quadratic attention complexity O(L 2 d) and bring about too high computational cost, especially when the sequence length L is much larger than the embedding dimension d. Therefore, the embodiment of the present disclosure can fuse features in a grouping manner to reduce computational cost. The electronic device 110 can determine multiple reference objects based on historical interaction information associated with the target object, and then divide the multiple reference objects into multiple groups of reference objects based on the historical interaction time corresponding to the multiple reference objects. Figure 3 , the user long sequence feature 302 includes multiple groups of reference objects. In this way, if a total of K groups are divided, the sequence length can be reduced to 1 / K of the original, thereby effectively performing space compression and achieving a trade-off between model efficiency and representation fidelity.

[0041] For example, for a common converter, the number of floating point operations is FLOPs. vanilla trans and parameters#Params vanilla trans It can be expressed as follows:

[0042] FLOPs vanilla trans =24Ld 2 +4L 2 d

[0043] #Params vanilla trans =12d 2 +13d

[0044] FLOPs Merge Token Indicates the number of floating-point operations corresponding to feature fusion, FLOPs vanilla It represents the number of floating-point operations before feature fusion. The ratio between the two is the attention complexity ratio as shown below:

[0045]

[0046] L represents the sequence length, d represents the embedding dimension, and K represents the number of groups. If L = 2048 and d = 32, the corresponding number of floating-point operations before feature fusion is approximately 587M, while the corresponding number of floating-point operations after feature fusion is approximately 336M, thus significantly improving computational efficiency.

[0047] In addition, feature fusion can not only reduce the computational complexity by shortening the sequence length, but also increase the number of parameters Θ merge , which not only improves the efficiency but also enhances the expressiveness of the model, thereby improving the overall performance of the model. The number of parameters after feature fusion Θ merge The calculation formula is as follows:

[0048] Θ merge =12K 2 d 2 +13Kd

[0049] In block 220 , the electronic device 110 updates, for each group of reference objects in the plurality of groups of reference objects, the feature portion corresponding to each group of reference objects in the first feature sequence based on the correlation between the reference objects in each group of reference objects to determine a second feature sequence.

[0050] In some embodiments, in order to better capture the time dynamics of the user's long sequence features, after obtaining the first feature sequence, refer to Figure 3 , a shared embedding layer 305 can be used to make multiple features in the first feature sequence share an embedding matrix to generate a vector representation, and then a position encoding layer 306 is used to encode the position information for each position in the sequence to update the first feature sequence.

[0051] In some embodiments, the position encoding layer 306 is implemented by allowing the model to learn the optimal representation of position during training. The position information encoded by the position encoding layer 306 can include two forms: an absolute time difference feature, which quantifies the time interval between historical interaction information and the target item; and a learnable absolute position embedding, which encodes the position of each feature in the sequence.

[0052] After position encoding, the obtained feature representation will be passed through a multi-layer perceptron to generate a feature sequence R:

[0053] R∈R (m+L)×d =[G∈R m×d ; H∈R L×d ]

[0054] G represents global features, and H represents user long sequence features. Next, the feature parts corresponding to each group of reference objects in the first feature sequence can be input into the feature fusion module 307 for feature update to determine a second feature sequence.

[0055] In some embodiments, simply connecting features within a group may result in insufficient interaction between features, thereby losing fine-grained details. Therefore, the electronic device 110 can determine, based on the feature portions corresponding to each group of reference objects, intra-group correlation features for each group of reference objects. The intra-group correlation features indicate the correlation between reference objects in each group of reference objects. Based on the intra-group correlation features, the feature portions corresponding to each group of reference objects in the first feature sequence are then updated to determine a second feature sequence.

[0056] For example, reference Figure 3 The first feature sequence includes feature portions 309 corresponding to each group of reference objects. Multiple adjacent features in each feature portion 309 can be interacted to obtain intra-group associated features 308. Then, the intra-group associated features 308 and the original feature portions 309 can be merged, and the merged features can be used to update the feature portions 309 corresponding to each group of reference objects in the first feature sequence to obtain a second feature sequence.

[0057] In some embodiments, the electronic device 110 may process the feature portion using a transformer unit to determine the intra-group correlation feature of each group of reference objects. For example, the intra-group correlation feature 308 is represented as follows:

[0058]

[0059] M i represents the intra-group correlation characteristics of group i, represents the embedding of the k-th reference object in the i-th group.

[0060] In this way, we can ensure that the interaction features within each group are effectively captured without the information loss caused by simply connecting the features within the group. And because the dimension and sequence length required to apply the transformer within the group are very small, in practical applications, applying the transformer within the group does not consume too many computational resources.

[0061] In block 230 , the electronic device 110 compresses a plurality of feature portions corresponding to the plurality of groups of reference objects in the second feature sequence to determine a third feature sequence, wherein a length of the third feature sequence is less than a length of the second feature sequence.

[0062] For example, reference Figure 3 , the second feature sequence is the input of the cross-attention module 308, and the third feature sequence is the output of the cross-attention module 308. The feature compression operation is performed by the cross-attention module 308, so the length of the third feature sequence is less than the length of the second feature sequence. The cross-attention mechanism is a mechanism for establishing a connection between two different sequences or feature sets, so that the model can focus on the relevant important information in the two sequences. The calculation formula of cross-attention Attention(Q,K,V) is as follows:

[0063] Q=OW Q , K=RW K ,V=RW V

[0064]

[0065] O is the query vector, R is the feature sequence, Q is the query matrix, K is the key matrix, V is the value matrix, M is the causal attention mask, W Q 、W K 、W V Indicates the shape is R d×d The query, key, and value projection matrices of .

[0066] In some embodiments, a causal attention mechanism can also be applied to improve model performance. The electronic device 110 can update multiple query vectors using a causal attention mask, which indicates that the attention information of the query vector is independent of the key vector with an index greater than the query vector. For example, Figure 3 , in the cross attention module 310, the query matrix generated in the previous step and the input sequence R∈R (m+L)×d , to apply the cross causal attention mechanism. The attention mask M is defined as:

[0067]

[0068] Applying causal masks can preserve the temporal correlation between sequence features on the one hand, and ensure the invisibility from sequence to candidate objects on the other hand, thus realizing a key-value cache service.

[0069] In some embodiments, reference Figure 3 Electronic device 110 may construct multiple query vectors based on the second feature sequence, where the multiple query vectors include at least one reference query vector 311 corresponding to multiple groups of reference objects, and the number of vectors in the at least one reference query vector 311 is less than the number of groups in the multiple groups of reference objects. The multiple query vectors are then updated based on the attention mechanism to determine a third feature sequence.

[0070] In some embodiments, the electronic device 110 can sample multiple embedding representations corresponding to multiple sampled reference objects in the multiple groups of reference objects from the second feature sequence. Then, based on the multiple embedding representations, a fourth feature sequence is constructed, and then multiple query vectors are constructed based on the fourth feature sequence. For example, according to a preset sampling strategy, k embedding representations H can be obtained by sampling multiple feature parts corresponding to multiple groups of reference objects in the second feature sequence. s ∈R k×d , then the part G∈R corresponding to the m global features in the second feature sequence m×d With k embedding representations H s ∈R k×d Perform splicing to obtain the query vector O, the formula is as follows:

[0071] O=[G;H s ]

[0072] In this way, attention can be focused on key local information and global contextual information, allowing the model to effectively capture specific sequential dependencies and broader contextual relationships.

[0073] In some embodiments, the preset sampling strategy may include selecting the k most recent embedding representations, uniformly sampling k embedding representations, initializing k learnable embedding representations, etc. Since there is a significant marginal effect between model performance and the number of sequence features, partial sampling can achieve a significant balance between computational overhead and model performance.

[0074] In some embodiments, after calculating the cross attention, the output can be passed to a feed-forward network for further processing.

[0075] In block 240 , the electronic device 110 determines a degree of association between the target object and the candidate objects based on the third feature sequence.

[0076] In some embodiments, the electronic device 110 may provide the target object with recommended content associated with the candidate object in response to the association degree being greater than a threshold value, referring to Figure 3 , a concatenation layer 313, a high-order MLP layer 314, and a prediction layer 315 can be applied to determine the degree of association between the target object and the candidate objects, such as click or conversion probability, as follows:

[0077] P(y=1|S u ,ud,v)∈[0,1]

[0078] y represents whether the target object u will interact with the target item v. The model can then be optimized using the binary cross entropy loss function, as follows:

[0079]

[0080] The historical interaction information can be expressed as D = {(S u ,ud,v,y)}, the model prediction probability can be expressed as

[0081] In some embodiments, the electronic device 110 may use at least one attention unit to update the third feature sequence. Then, based on the updated third feature sequence, the degree of association between the target object and the candidate object is determined. For example, referring to Figure 3 After the cross-attention module 310, a self-attention module 312 is included. The self-attention module 312 includes multiple self-attention layers. These self-attention layers are used to learn and obtain the internal relationships in the feature sequence, allowing the model to capture the dependencies and patterns between features in the feature sequence. Each self-attention layer is followed by a feedforward layer to further process the information learned by the self-attention mechanism. The calculation formula of the self-attention mechanism is as follows:

[0082]

[0083] The query matrix Q, key matrix K and value matrix V are obtained by applying the linear projection matrix W to the output of the previous layer. Q 、W K and W V Got it.

[0084] Since the self-attention module 312 includes multiple self-attention layers, the representation of the input sequence can be iteratively optimized through these self-attention layers. After passing through these self-attention layers, a compressed output representation is generated. This output representation is the final output result of the attention mechanism, that is, the third feature sequence, which will be used for downstream prediction tasks. The conversion process can be as follows:

[0085] CrossAttn(O,R)→SelfAttn(·)×N

[0086] CrossAttn stands for the cross attention mechanism, and SelfAttn stands for the self-attention mechanism. Through the cross attention mechanism and self-attention mechanism, the model can effectively process long feature sequences while utilizing global context information and internal dependencies.

[0087] In some embodiments, a causal attention mechanism may also be applied in the self-attention module 312, and the electronic device 110 may update multiple query vectors using a causal attention mask, where the causal attention mask indicates that the attention information of the query vector is independent of the key vectors whose index is greater than the query vector. For example, Figure 3,In the self-attention module 312, a cross causal attention mechanism can be applied by using the query matrix and input sequence generated in the previous step.

[0088] In some embodiments, to improve the efficiency of reasoning when scoring candidate objects, a key-value caching mechanism can be applied. The electronic device 110 can cache first attention information associated with multiple groups of reference objects, and then use the cached first attention information to determine second attention information associated with the second candidate object. Then, based on the first attention information and the second attention information, a second degree of association between the target object and the second candidate object is determined. The candidate object is the first candidate object, and the degree of association is the first degree of association.

[0089] Since the user long sequence feature remains unchanged in each candidate object, its internal representation only needs to be calculated once and can be reused. Therefore, the attention calculation between the part of the feature sequence related to the user long sequence feature and the part related to the global feature can be decoupled. For example, the feature sequence can be divided into the part related to the user long sequence feature and the part related to the global feature. Figure 5 , the white circle represents the part related to the user's long sequence features, and the black circle represents the part related to the global features. Then refer to Figure 5 In step 501, the first attention information associated with the part related to the user's long sequence features is pre-calculated and cached, that is, the projection matrix of the key and value related to the user's long sequence features will be pre-calculated and cached. Then refer to Figure 5 In step 502, the cached first attention information is used to determine the second attention information associated with the second candidate object. This method avoids redundant calculations and significantly reduces service latency. This attention calculation method can be applied to both the cross-attention module and the self-attention module.

[0090] In some embodiments, reference Figure 4 During the model training process, training data 401 can be obtained through batch processing or streaming processing. Batch processing is to import large batches of data into the system at one time on a regular basis, while streaming processing is to continuously ingest data into the system in real time or near real time, and each record is processed immediately after it arrives. The acquired data is then preprocessed by the data stream module 402. The processed training data can then be distributed to the execution module 403, which includes multiple graphics processing units (GPUs). During the training process, dense parameters and sparse parameters are calculated and updated together in the same training step, without the need for an external parameter server component, so that it can be efficiently expanded across devices and multiple nodes, thereby supporting the training and inference of large-scale parameter models.

[0091] In some embodiments, mixed precision training can be used to reduce the computational overhead caused by scaling dense models. Users can configure precision at the model level, applying higher precision to key components and lower precision elsewhere. In addition, to alleviate the memory pressure on the GPU during training, a recalculation strategy can be used in conjunction with mixed precision training. Reverse mode automatic differentiation can be used for gradient calculations, which is more efficient than forward mode automatic differentiation because only one backpropagation is required to calculate the partial derivatives of all outputs with respect to all inputs, while forward mode automatic differentiation may require multiple forward propagations to complete the same task, but reverse mode automatic differentiation needs to store all intermediate activation values ​​during the forward propagation.

[0092] Since these intermediate activation values ​​can become a major memory bottleneck, it is possible to support recomputation declarations during model construction, so that selected activation values ​​can be discarded during the forward propagation and recomputed during the backward propagation. This approach saves memory usage at the expense of increased computation.

[0093] The disclosed embodiments can promote information fusion across the entire sequence and stabilize attention distribution by introducing global features. Feature fusion reduces computational complexity, and by generating intra-group correlation features, computational costs are effectively reduced while maintaining intra-group dependencies. Furthermore, a hybrid attention mechanism, including cross-attention and self-attention, is employed to improve computational efficiency while maintaining performance.

[0094] Example devices and equipment

[0095] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 6 FIG2 is a schematic block diagram of an apparatus 600 for information processing according to certain embodiments of the present disclosure. The apparatus 600 may be implemented as or included in the electronic device 110 discussed above. Each module / component in the apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.

[0096] like Figure 6As shown, the device 600 includes a construction module 610, which is configured to construct a first feature sequence based on the target object, candidate objects and multiple groups of reference objects associated with the target object; an update module 620, which is configured to update the feature parts corresponding to each group of reference objects in the first feature sequence based on the correlation between the reference objects in each group of reference objects to determine a second feature sequence; a compression module 630, which is configured to compress multiple feature parts corresponding to the multiple groups of reference objects in the second feature sequence to determine a third feature sequence, wherein the length of the third feature sequence is less than the length of the second feature sequence; and a determination module 640, which is configured to determine the degree of association between the target object and the candidate object based on the third feature sequence.

[0097] In some embodiments, the apparatus 600 is further configured to determine a plurality of reference objects based on historical interaction information associated with the target object; and divide the plurality of reference objects into a plurality of groups of reference objects based on historical interaction times corresponding to the plurality of reference objects.

[0098] In some embodiments, the update module 620 is configured to determine the intra-group association features of each group of reference objects based on the feature parts corresponding to each group of reference objects, where the intra-group association features indicate the correlation between the reference objects in each group of reference objects; and to update the feature parts corresponding to each group of reference objects in the first feature sequence based on the intra-group interaction features to determine the second feature sequence.

[0099] In some embodiments, the updating module 620 is further configured to process the feature portion using a converter unit to determine intra-group associated features of each group of reference objects.

[0100] In some embodiments, the compression module 630 is configured to construct multiple query vectors based on the second feature sequence, the multiple query vectors including at least one reference query vector corresponding to multiple groups of reference objects, and the number of vectors of at least one reference query vector is less than the number of groups of the multiple groups of reference objects; and update the multiple query vectors based on the attention mechanism to determine a third feature sequence.

[0101] In some embodiments, the compression module 630 is further configured to sample, from the second feature sequence, multiple embedding representations corresponding to multiple sampled reference objects in multiple groups of reference objects; construct a fourth feature sequence based on the multiple embedding representations; and construct multiple query vectors based on the fourth feature sequence.

[0102] In some embodiments, the compression module 630 is further configured to update the plurality of query vectors using a causal attention mask indicating that attention information of the query vector is independent of key vectors having an index greater than the query vector.

[0103] In some embodiments, the determination module 640 is configured to update the third feature sequence using at least one attention unit; and determine the degree of association between the target object and the candidate object based on the updated third feature sequence.

[0104] In some embodiments, the device 600 is also configured to cache first attention information associated with multiple groups of reference objects; and determine second attention information associated with the second candidate object using the cached first attention information; and determine a second degree of association between the target object and the second candidate object based on the first attention information and the second attention information.

[0105] In some embodiments, the apparatus 600 is further configured to provide the target object with recommended content associated with the candidate object in response to the association degree being greater than a threshold.

[0106] The units included in the device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 600 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0107] Figure 7 1 shows a block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented. Figure 7 The illustrated electronic device 700 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 7 The electronic device 700 shown can be used to implement Figure 1 An electronic device 110 is shown.

[0108] like Figure 7 As shown, electronic device 700 is in the form of a general electronic device. Components of electronic device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processing unit 710 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of electronic device 700.

[0109] The electronic device 700 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the electronic device 700.

[0110] The electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 7 As shown in FIG, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules configured to perform the various methods or actions of various embodiments of the present disclosure.

[0111] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 700 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 700 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0112] Input device 750 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 760 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 700 may also communicate with one or more external devices (not shown) via communication unit 740 as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with electronic device 700, or with any device that allows electronic device 700 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0113] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0114] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0115] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0116] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0117] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0118] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for information processing, comprising: constructing a first feature sequence based on a target object, candidate objects, and a plurality of groups of reference objects associated with the target object; For each group of reference objects in the plurality of groups of reference objects, based on correlations between reference objects in each group of reference objects, updating a feature portion corresponding to each group of reference objects in the first feature sequence to determine a second feature sequence; compressing a plurality of feature portions corresponding to the plurality of groups of reference objects in the second feature sequence to determine a third feature sequence, wherein a length of the third feature sequence is less than a length of the second feature sequence; as well as Based on the third feature sequence, a degree of association between the target object and the candidate objects is determined.

2. The method according to claim 1, further comprising: determining a plurality of reference objects based on historical interaction information associated with the target object; as well as The multiple reference objects are divided into the multiple groups of reference objects based on historical interaction times corresponding to the multiple reference objects.

3. The method according to claim 1, wherein The updating, for each group of reference objects in the plurality of groups of reference objects, of the feature portion corresponding to each group of reference objects in the first feature sequence based on correlations between the reference objects in the groups of reference objects to determine a second feature sequence, includes: determining, based on the feature parts corresponding to the groups of reference objects, intra-group correlation features of the groups of reference objects, the intra-group correlation features indicating the correlations between the reference objects in the groups of reference objects; and Based on the intra-group associated features, the feature parts corresponding to the groups of reference objects in the first feature sequence are updated to determine a second feature sequence.

4. The method according to claim 3, wherein: The determining, based on the feature parts corresponding to the groups of reference objects, the intra-group association features of the groups of reference objects includes: The feature portions are processed using a converter unit to determine intra-group correlation features for each group of reference objects.

5. The method according to claim 1, wherein The compressing the plurality of feature parts corresponding to the plurality of groups of reference objects in the second feature sequence to determine a third feature sequence includes: constructing a plurality of query vectors based on the second feature sequence, wherein the plurality of query vectors includes at least one reference query vector corresponding to the plurality of groups of reference objects, and the number of vectors in the at least one reference query vector is less than the number of groups of the plurality of groups of reference objects; and Based on an attention mechanism, the multiple query vectors are updated to determine the third feature sequence.

6. The method according to claim 5, wherein: The constructing of multiple query vectors based on the second feature sequence includes: sampling, from the second feature sequence, a plurality of embedding representations corresponding to a plurality of sampled reference objects in the plurality of groups of reference objects; constructing a fourth feature sequence based on the multiple embedding representations; and The multiple query vectors are constructed based on the fourth feature sequence.

7. The method according to claim 5, wherein: The updating of the plurality of query vectors based on the attention mechanism includes: The plurality of query vectors are updated using a causal attention mask that indicates that attention information of a query vector is independent of key vectors having an index greater than that of the query vector.

8. The method according to claim 1, wherein The determining, based on the third feature sequence, the degree of association between the target object and the candidate object includes: Updating the third feature sequence using at least one attention unit; and The degree of association between the target object and the candidate object is determined based on the updated third feature sequence.

9. The method according to claim 1, wherein The candidate object is a first candidate object, the association degree is a first association degree, and the method further includes: caching first attention information associated with the plurality of groups of reference objects; and Determining second attention information associated with a second candidate object using the cached first attention information; and Based on the first attention information and the second attention information, a second degree of association between the target object and a second candidate object is determined.

10. The method according to claim 1, further comprising: In response to the association degree being greater than a threshold, recommended content associated with the candidate object is provided to the target object.

11. An information processing device, comprising: a construction module configured to construct a first feature sequence based on a target object, candidate objects, and a plurality of groups of reference objects associated with the target object; an updating module configured to update, for each group of reference objects in the plurality of groups of reference objects, a feature portion corresponding to each group of reference objects in the first feature sequence based on correlations between reference objects in the groups of reference objects, so as to determine a second feature sequence; a compression module configured to compress a plurality of feature portions corresponding to the plurality of groups of reference objects in the second feature sequence to determine a third feature sequence, wherein a length of the third feature sequence is less than a length of the second feature sequence; as well as The determination module is configured to determine the degree of association between the target object and the candidate object based on the third feature sequence.

12. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.

13. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 10 when executed by a processor.

14. A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.