News material recommendation method and device based on angle decoupling and cross-modal matching and electronic equipment

By building structured instruction templates and dynamic weight optimization models, analyzing news texts and generating matching weights, the problems of insufficient multimodal semantic coupling and dynamic scene adaptation in traditional methods are solved, and high-precision news material recommendations are achieved.

CN120561374APending Publication Date: 2025-08-29WUHAN UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510654524.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Traditional news material recommendation methods cannot effectively distinguish information coupling in different semantic dimensions, and cannot adapt to the dynamic changes in semantic weights during news dissemination, resulting in multimodal semantic coupling, lack of angle-specific weights and insufficient adaptation of dynamic scenes.

Method used

Build a structured instruction template, analyze news text through a big model, extract multi-angle text elements, and perform cross-modal encoding, generate text vectors and image vectors, use dynamic weight optimization model to generate matching weights, and calculate the similarity between compound semantic vectors and image vectors for recommendation.

Benefits of technology

It realizes high-precision picture and text matching, improves the accuracy of news material recommendations, can adapt to the scene needs of different news types, adapt to the differentiated needs of news types, and improves the efficiency and accuracy of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561374A_ABST
    Figure CN120561374A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of news media, and provides a news material recommendation method and device based on angle decoupling and cross-modal matching and electronic equipment. Comprising the following steps: constructing a structured instruction template; calling the large model to perform semantic analysis on the input text based on the structured instruction template, and extracting multi-angle text elements; respectively carrying out cross-modal coding on the multi-angle text elements and the image materials to generate text vectors and image vectors; based on the text vector and the image vector, generating a matching weight of each text element through a dynamic weight optimization model; and generating a composite semantic vector according to the matching weight, calculating the similarity between the composite semantic vector and the image vector, and recommending matching materials based on similarity sorting. A structured text feature decoupling model and a Bayesian weight learning framework are constructed, so that the model can capture associated elements of text angles such as a title and a main body and visual features, and high-precision image-text matching and recommendation accuracy improvement in a complex scene are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of news media technology, and specifically to a news material recommendation method, device, and electronic device based on angle decoupling and cross-modal matching. Background Art

[0002] With the explosive growth of media content and the refinement of user needs, news material recommendation faces the dual challenges of multimodal semantic alignment and dynamic scene adaptation. Traditional methods rely on keyword matching or fixed-weight feature fusion, which makes it difficult to cope with the multi-angle semantic complexity of news content: the different parts of a news text, such as the title, body, and tags, carry differentiated semantic functions. For example, the title usually focuses on the core elements of the event, while the body contains complete background information and a logical chain. Existing models often treat them as single-modal input, resulting in information coupling between different semantic dimensions. Although cross-modal models such as CLIP can achieve semantic alignment of images and text, their global feature averaging processing mechanism cannot effectively distinguish the contribution weights of different text angles to visual matching, and the static mapping relationship is difficult to adapt to the dynamic changes in semantic weights during news dissemination. For example, current affairs news needs to strengthen the weight of timeliness icons, while in-depth reports need to focus on the long-term relevance of background information. The fixed weight strategy of traditional methods is obviously difficult to meet such scenario-based needs.

[0003] To solve the problem of multimodal angle decoupling, traditional methods attempt to combine text attention mechanisms with visual feature pyramids, but they still have the following drawbacks:

[0004] Inadequate multimodal semantic alignment accuracy: Traditional methods rely on keyword matching or fixed-weight feature fusion, making it difficult to handle the multi-angle semantic complexity of news content. Different parts of a news text, such as the title and body, have distinct semantic functions. Existing models treat these as a single modal input, resulting in information coupling and an inability to effectively distinguish between different semantic dimensions. While cross-modal models such as CLIP can achieve semantic alignment between images and text, their global feature averaging mechanism cannot accurately distinguish the contribution weights of different text angles to visual matching.

[0005] Difficulty adapting to dynamic scenarios: Traditional methods' fixed weighting strategies cannot accommodate the dynamic changes in semantic weights during news dissemination. For example, current affairs news and in-depth reports place different emphasis on different information. Traditional methods struggle to adaptively adjust the matching priority of different text angles based on the differences in news types, making them ineffective in addressing complex scenarios.

[0006] In summary, it is necessary to solve the core problems of multimodal semantic coupling, lack of angle-specific weights, and insufficient adaptation to dynamic scenes in traditional news material recommendation. Summary of the Invention

[0007] In view of this, the embodiments of the present application provide a news material recommendation method, device and electronic device based on angle decoupling and cross-modal matching, which solves the core problems of multimodal semantic coupling, lack of angle-specific weights and insufficient dynamic scene adaptation in traditional news material recommendation.

[0008] A first aspect of an embodiment of the present application provides a news material recommendation method based on angle decoupling and cross-modal matching, comprising:

[0009] Build structured instruction templates;

[0010] Based on the structured instruction template, the large model is called to perform semantic analysis on the input text and extract multi-angle text elements;

[0011] Cross-modal encoding is performed on the multi-angle text elements and image materials to generate text vectors and image vectors;

[0012] Based on the text vector and the image vector, generating a matching weight for each text element through a dynamic weight optimization model;

[0013] A composite semantic vector is generated according to the matching weight, a similarity between the composite semantic vector and the image vector is calculated, and matching materials are recommended based on the similarity ranking.

[0014] A second aspect of an embodiment of the present application provides a news material recommendation device based on angle decoupling and cross-modal matching, including:

[0015] Structured instruction building module, used to build structured instruction templates;

[0016] A large model text parsing module is used to call the large model based on the structured instruction template to perform semantic parsing on the input text and extract multi-angle text elements;

[0017] A feature encoding module, configured to perform cross-modal encoding on the multi-angle text elements and image materials to generate text vectors and image vectors;

[0018] A weight optimization module, configured to generate a matching weight for each text element based on the text vector and the image vector through a dynamic weight optimization model;

[0019] The matching recommendation module is used to generate a composite semantic vector according to the matching weight, calculate the similarity between the composite semantic vector and the image vector, and recommend matching materials based on the similarity ranking.

[0020] A third aspect of an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the news material recommendation method based on angle decoupling and cross-modal matching as provided in the first aspect of the embodiment of the present application.

[0021] A fourth aspect of the embodiments of the present application provides a computer program product, including a computer program. When the computer program is executed, the method according to the first aspect of the embodiments of the present application is executed.

[0022] The first aspect of the embodiment of the present application provides a news material recommendation method based on angle decoupling and cross-modal matching, which is achieved by constructing a structured instruction template; calling a large model based on the structured instruction template to perform semantic analysis on the input text and extract multi-angle text elements; cross-modally encoding the multi-angle text elements and image materials to generate text vectors and image vectors; based on the text vectors and image vectors, generating matching weights for each text element through a dynamic weight optimization model; generating a composite semantic vector based on the matching weights, calculating the similarity between the composite semantic vector and the image vector, and recommending matching materials based on the similarity ranking. A structured text feature decoupling model and a Bayesian weight learning framework are constructed, breaking through the limitation of traditional algorithms that regard text as a single modal input, enabling the model to capture the correlation elements between text angles and visual features such as titles and text, achieving high-precision image-text matching and improving recommendation accuracy in complex scenarios. The core problems of multi-modal semantic coupling, lack of angle-specific weights, and insufficient dynamic scene adaptation in traditional news material recommendation are solved.

[0023] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0025] Figure 1 This is a flow chart of a news material recommendation method based on angle decoupling and cross-modal matching provided by an embodiment of the present application;

[0026] Figure 2 2 is a structural diagram of a news material recommendation device based on angle decoupling and cross-modal matching provided by an embodiment of the present application;

[0027] Figure 3 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0029] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0030] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0031] like Figure 1 As shown, the news material recommendation method based on angle decoupling and cross-modal matching provided by the embodiment of the present application includes the following steps S101 to S106:

[0032] Step S101: A news material recommendation method based on angle decoupling and cross-modal matching, characterized by comprising:

[0033] Step S102: constructing a structured instruction template;

[0034] Step S103: Based on the structured instruction template, the large model is called to perform semantic analysis on the input text to extract multi-angle text elements;

[0035] Step S104: cross-modal encoding is performed on the multi-angle text elements and image materials to generate text vectors and image vectors;

[0036] Step S105: generating a matching weight for each text element through a dynamic weight optimization model based on the text vector and the image vector;

[0037] Step S106 : generating a composite semantic vector according to the matching weight, calculating the similarity between the composite semantic vector and the image vector, and recommending matching materials based on the similarity ranking.

[0038] The embodiment of the present application guides the large model to conduct multi-angle semantic analysis of news texts by constructing a structured instruction template, breaking through the limitation of traditional methods that regard text as a single semantic unit. Based on the text vectors and image vectors generated by cross-modal coding, the dynamic weight optimization model is combined to realize element-level matching weight tuning, and finally the recommendation is completed by calculating the similarity between the composite semantic vector and the image vector. The structured instruction template injects professional knowledge through domain rules to solve the problem of fuzzy analysis of traditional large models in news scenarios; the multi-angle element decoupling coding retains differentiated semantic information such as titles and core events, and provides fine-grained input for dynamic weight optimization; the dynamic weight model replaces the fixed weight strategy, so that the recommendation system can adapt to the semantic matching needs of different scenarios such as current affairs news and in-depth reports, and significantly improve the accuracy of cross-modal alignment.

[0039] In one embodiment, constructing a structured instruction template includes:

[0040] Integrate news text, task objective description and domain mapping rules into predefined templates to generate structured instruction input templates.

[0041] In the application, building a structured instruction template means building a structured Prompt input template to guide the large model to parse the news text through structured semantic extraction instructions.

[0042] The structured instruction template constructed in the embodiment of the present application mainly includes the following three parts:

[0043] 1) News text: that is, complete news text data including title and text.

[0044] 2) Task objective description: Clearly inform the large model of the parsing task to be performed, such as "extracting the core elements of a news event", and specify the JSON format output requirements.

[0045] 3) Domain feature mapping rules: Guide the large model through examples to understand the semantic associations unique to the news domain, including visual metaphor mapping and event structure constraints.

[0046] Finally, the three parts are integrated into the predefined Prompt template to generate a complete Prompt string, thereby converting the news text into a structured Prompt input to facilitate the subsequent large model to parse the news text.

[0047] The embodiment of the present application defines the specific construction method of the structured instruction template, integrating news text, task objectives and domain mapping rules into an instruction input template. The task objective description clearly defines the parsing element type (such as core events, visual metaphors), avoiding the omission of elements caused by vague instructions in traditional methods; the domain mapping rules explicitly associate news types with visual symbols (for example, current affairs news maps government badge icons), guiding the large model to generate symbolic features that can be aligned across modalities, solving the problem of the semantic gap between text and images; predefined templates force the output of structured data (such as JSON format), ensuring that downstream modules can directly call the parsing results, eliminating data processing redundancy caused by unstructured output.

[0048] In one embodiment, the multi-angle text element includes:

[0049] Core event elements, significant entity elements, key action elements, emotional tendency elements, scene feature elements, visual metaphor elements and important data elements.

[0050] In the application, the large model called by Prompt performs deep semantic analysis of news texts, extracting seven major elements in parallel: core events, salient entities, key actions, emotional tendencies, scene features, visual metaphors, and important data. Specifically, the large model text parsing module uses the large model called by Prompt to perform deep semantic analysis of news texts. By loading the pre-trained large model and injecting domain adaptation parameters, the model first performs noise filtering and semantic disambiguation on the input news text. Then, based on the structured instructions defined in Prompt, it extracts seven major elements in parallel: core events, salient entities, key actions, emotional tendencies, scene features, visual metaphors, and important data. Among them, the analysis of core events uses a joint model of dependency parsing and semantic role labeling to accurately locate the "person-action-object" triples; the mapping of visual metaphors combines the attention mechanism to dynamically focus on keywords to generate a list of candidate visual symbols. The parsed results are forced to be output in JSON format.

[0051] The embodiment of the present application clarifies that multi-angle text elements include seven types of elements, such as core events and significant entities. The text semantics are refined and decoupled through the element classification system. The title focuses on the core of the event and the text analyzes the background logic, solving the matching deviation caused by the semantic coupling of the traditional model; the core event elements are represented by the person-action-object triplet, accurately capturing the logical structure of the news event and enhancing the correlation between the event elements and the visual elements; the visual metaphor elements map abstract semantics (such as "economic recession") into specific symbols (such as "downward arrow"), establishing an explicit association channel between the deep semantics of the text and the visual features, and improving the interpretability of cross-modal matching.

[0052] In one embodiment, cross-modal encoding of the multi-angle text elements and image materials to generate text vectors and image vectors includes:

[0053] Input the title and the multi-angle text elements as independent semantic units into the CLIP text encoder to generate independent text vectors;

[0054] Input all the image materials in the material library into the CLIP visual encoder to generate a global image vector;

[0055] Dimensionality reduction processing is performed on the independent text vector and the global image vector respectively to obtain the text vector and the image vector.

[0056] In the application, eight types of text features including titles, seven major elements and image features are encoded separately through the CLIP model to generate semantic vectors of unified dimensions, and a cross-modal retrieval library is constructed to achieve alignment and unified representation of multimodal features. Specifically, the feature encoding module maps the parsed text features and image features to a unified semantic space. For the text side, the title and seven major elements are input into the CLIP text encoder as independent semantic units to generate 8 512-dimensional vectors, which are then reduced to 256 dimensions through a linear transformation layer with shared weights, retaining key semantic information and reducing computational redundancy. On the image side, all images in the material library are input into the CLIP visual encoder to generate a 512-dimensional vector, and the spatial dimension is compressed through mean pooling to generate a global feature vector with translation invariance.

[0057] The embodiment of the present application limits the specific implementation of the cross-modal encoding process, inputs each text element as an independent unit into the CLIP encoder, avoids the semantic interference caused by feature fusion, and retains the differentiated semantics of different angles such as the title and text; performs dimensionality reduction processing on the text vector to reduce the computational complexity while retaining key semantic information, and solves the storage and matching efficiency problems brought by high-dimensional vectors; image encoding uses mean pooling to compress spatial dimensions, generate global features with translation invariance, enhance the representation ability of the main content of the image, and reduce the interference of background noise on the matching results.

[0058] In one embodiment, generating the matching weight of each text element by a dynamic weight optimization model based on the text vector and the image vector includes:

[0059] Define the constrained weight search space;

[0060] In descending order of similarity ranking, multiple image vectors are selected from the maximized candidate image library, and the average cosine similarity of multiple image vectors is used as the objective function. The tree-structured Parson estimator sampling algorithm is used to perform iterative optimization of weighted combination until the iteration termination condition is reached, and the matching weight of each text element is output.

[0061] In the application, the matching weights of eight categories of text features are Bayesian optimized using the Optuna framework. The weight combinations of each dimension are dynamically adjusted, and iterative optimization is performed with the goal of maximizing image-text matching similarity, ultimately outputting the optimal weight parameters. Specifically, the weight optimization module dynamically tunes the weights of the eight categories of text features using the Optuna framework. First, a constrained search space is defined: a hard constraint requires that the weights sum to 1, implicitly normalized using the Softmax function; a soft constraint introduces a feature correlation matrix (e.g., a negative correlation between the weight of the title and the core event, and a positive correlation between the weight of sentiment and the weight of the visual metaphor), and a penalty term is added to the objective function to guide the search direction. The optimization goal is set to maximize the average cosine similarity of the top five images. The TPE (Tree-structured Parzen Estimator) sampling algorithm is used to prioritize high-potential areas. Each iteration randomly generates weight combinations and calculates the target value. To accelerate convergence, a dynamic early stopping strategy is designed: when the objective function improvement is less than 0.1% for 20 consecutive iterations, the optimization is considered converged and the search is terminated, ultimately outputting the optimal weight parameters.

[0062] The embodiment of the present application proposes a dynamic weight optimization mechanism based on a constrained search space and a TPE algorithm, with the average value of Top-N image similarities as the optimization target, so that weight tuning focuses on highly relevant candidate materials and avoids low-quality matching samples from interfering with the optimization direction; the TPE sampling algorithm predicts high-potential weight areas by constructing a probability model, significantly improving the convergence speed compared to random search or grid search; the iterative termination condition dynamically determines the optimized convergence, avoiding invalid calculations while ensuring accuracy, and is particularly suitable for real-time recommendation scenarios with large-scale updates of news material libraries.

[0063] In one embodiment, the hard constraints of the constrained weight search space are normalized by the Softmax function so that the sum of each weight is 1, and the soft constraints of the constrained weight search space impose penalty terms on the weight combination based on the feature correlation matrix, and the feature correlation matrix includes a weight negative correlation matrix between the title and the core event elements and a positive correlation matrix between the emotional tendency elements and the visual metaphor elements.

[0064] The embodiment of the present application ensures the rationality of the probability distribution of weight parameters through Softmax hard constraints, preventing matching deviations caused by excessive weight of a single factor; the feature correlation matrix introduces domain knowledge (such as the negative correlation between titles and core events), and guides the weight combination to conform to the semantic laws of news through penalty items, thereby solving the overfitting risk of pure data-driven models; positive / negative correlation constraints encode differentiated matching strategies for news types (for example, current affairs news strengthens the weight of core events), so that the optimization process has the dual advantages of data-driven and knowledge-guided.

[0065] In one embodiment, generating a composite semantic vector according to the matching weight includes:

[0066] The text vectors are linearly superimposed according to weights to generate a composite semantic vector consistent with the dimension of the image vector.

[0067] The linear superposition in the embodiment of the present application is mathematically equivalent to the weighted fusion of multi-angle semantics, which not only retains the contribution information of each factor, but also meets the computability requirements of the vector space; the dimensional consistency design ensures that the similarity of text and image vectors can be directly calculated, eliminating the projection error caused by the mismatch of cross-modal vector dimensions in traditional methods; the weight parameters dynamically act on the vector superposition process, so that the recommendation results can reflect the changes in semantic weights brought about by the switching of news hotspots in real time (for example, the increase in the weight of the time factor in emergencies).

[0068] In one embodiment, calculating the similarity between the composite semantic vector and the image vector, and recommending matching materials based on the similarity ranking, includes:

[0069] Calculate the cosine similarity between the composite semantic vector and the image vector, and recommend matching materials in descending order of the cosine similarity.

[0070] In the app, the CLIP model is used to calculate image-text similarity, and matching materials are recommended based on this similarity. The cosine similarity metric effectively captures directional consistency in vectors and is more suitable for semantic matching of high-dimensional sparse vectors than the Euclidean distance. A descending sorting mechanism ensures that recommendations are presented in descending order of matching, aligning with user browsing habits and system retrieval efficiency requirements.

[0071] This method adopts a Bayesian dynamic weight optimization framework to achieve adaptive fusion and scenario-based adaptation of cross-modal features. In response to the problem that traditional fixed weight strategies cannot adapt to the differentiated needs of news types, the present invention proposes a constrained Bayesian search mechanism, which dynamically learns the optimal weight combination in different scenarios through the joint optimization of hard constraints, namely weight normalization, and soft constraints, namely feature correlation matrices. For example, current affairs news automatically strengthens the weight of the event subject, and social news focuses on emotional tendencies and visual metaphor associations. This mechanism enables the system to show significant advantages in both the rapid matching of breaking news and the precise matching of in-depth reports. At the same time, combined with the TPE sampling algorithm and the dynamic early stopping strategy, it greatly shortens the optimization time and achieves a double breakthrough in efficiency and accuracy.

[0072] This method provides cross-domain migration capabilities, supporting rapid adaptation to different vertical news scenarios. By adjusting the domain feature mapping rules in Prompt, for example, changing "financial policy → stock chart" to "medical policy → human anatomy diagram," the system can seamlessly migrate to material recommendation scenarios in fields such as healthcare and education.

[0073] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0074] The present application also provides a news material recommendation device based on angle decoupling and cross-modal matching, which is configured to perform the steps described in the aforementioned method for recommending news material based on angle decoupling and cross-modal matching. The news material recommendation device based on angle decoupling and cross-modal matching can be a virtual appliance within an electronic device, executed by a processor within the electronic device, or it can be the electronic device itself.

[0075] like Figure 2 As shown, the news material recommendation device 100 based on angle decoupling and cross-modal matching provided in an embodiment of the present application includes:

[0076] The structured instruction construction module 101 is used to construct a structured instruction template;

[0077] A large model text parsing module 102 is used to call the large model based on the structured instruction template to perform semantic parsing on the input text and extract multi-angle text elements;

[0078] A feature encoding module 103 is configured to perform cross-modal encoding on the multi-angle text elements and image materials to generate text vectors and image vectors;

[0079] A weight optimization module 104 is configured to generate a matching weight for each text element based on the text vector and the image vector using a dynamic weight optimization model;

[0080] The matching recommendation module 105 is configured to generate a composite semantic vector according to the matching weight, calculate the similarity between the composite semantic vector and the image vector, and recommend matching materials based on the similarity ranking.

[0081] In application, each module in the news material recommendation device based on angle decoupling and cross-modal matching can be a software program module, or can be implemented through different logic circuits integrated in a processor, or can be implemented through multiple distributed processors.

[0082] like Figure 3 As shown, the embodiment of the present application further provides an electronic device 200, including: at least one processor 201 ( Figure 3 Only one processor is shown in the figure), a memory 202, and a computer program 203 stored in the memory 202 and executable on at least one processor 201. When the processor 201 executes the computer program 203, the steps in the above-mentioned various method embodiments are implemented.

[0083] In applications, electronic devices may include, but are not limited to, processors and memories. Those skilled in the art will appreciate that Figure 3 The electronic device is merely an example and does not limit the electronic device. The electronic device may include more or fewer components than shown in the figure, or may include a combination of certain components or different components.

[0084] In applications, the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0085] In applications, in some embodiments, the memory can be an internal storage unit of an electronic device, such as a hard disk or memory of the electronic device. In other embodiments, the memory can also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card, etc. Furthermore, the memory can also include both an internal storage unit of the electronic device and an external storage device. The memory is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of a computer program. The memory can also be used to temporarily store data that has been output or is about to be output.

[0086] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0087] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0088] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0089] An embodiment of the present application provides a computer program product, including a computer program. When the computer program product runs on an electronic device, the electronic device can implement the steps in the above-mentioned various method embodiments when executing the computer program product.

[0090] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, a computer-readable medium cannot be an electric carrier signal or a telecommunication signal.

[0091] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0092] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0093] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0094] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0095] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A news material recommendation method based on angle decoupling and cross-modal matching, characterized in that: include: Build structured instruction templates; Based on the structured instruction template, the large model is called to perform semantic analysis on the input text and extract multi-angle text elements; Cross-modal encoding is performed on the multi-angle text elements and image materials to generate text vectors and image vectors; Based on the text vector and the image vector, generating a matching weight for each text element through a dynamic weight optimization model; A composite semantic vector is generated according to the matching weight, a similarity between the composite semantic vector and the image vector is calculated, and matching materials are recommended based on the similarity ranking.

2. The news material recommendation method based on angle decoupling and cross-modal matching according to claim 1, characterized in that: The constructing of the structured instruction template includes: Integrate news text, task objective description and domain mapping rules into predefined templates to generate structured instruction input templates.

3. The news material recommendation method based on angle decoupling and cross-modal matching according to claim 1, characterized in that: The multi-angle text elements include: Core event elements, significant entity elements, key action elements, emotional tendency elements, scene feature elements, visual metaphor elements and important data elements.

4. The news material recommendation method based on angle decoupling and cross-modal matching according to claim 1, characterized in that: The cross-modal encoding of the multi-angle text elements and image materials to generate text vectors and image vectors includes: Input the title and the multi-angle text elements as independent semantic units into the CLIP text encoder to generate independent text vectors; Input all the image materials in the material library into the CLIP visual encoder to generate a global image vector; Dimensionality reduction processing is performed on the independent text vector and the global image vector respectively to obtain the text vector and the image vector.

5. The news material recommendation method based on angle decoupling and cross-modal matching according to claim 1, characterized in that: Generating the matching weight of each text element by a dynamic weight optimization model based on the text vector and the image vector includes: Define the constrained weight search space; In descending order of similarity ranking, multiple image vectors are selected from the maximized candidate image library, and the average cosine similarity of multiple image vectors is used as the objective function. The tree-structured Parson estimator sampling algorithm is used to perform iterative optimization of weighted combination until the iteration termination condition is reached, and the matching weight of each text element is output.

6. The news material recommendation method based on angle decoupling and cross-modal matching according to claim 5, characterized in that: The hard constraints of the constrained weight search space are normalized by the Softmax function so that the sum of each weight is 1. The soft constraints of the constrained weight search space impose penalty terms on the weight combination based on the feature correlation matrix. The feature correlation matrix includes the weight negative correlation matrix of the title and the core event elements and the positive correlation matrix of the emotional tendency elements and the visual metaphor elements.

7. The news material recommendation method based on angle decoupling and cross-modal matching according to claim 1, characterized in that: Generating a composite semantic vector according to the matching weight includes: The text vectors are linearly superimposed according to weights to generate a composite semantic vector consistent with the dimension of the image vector.

8. The news material recommendation method based on angle decoupling and cross-modal matching according to claim 1, characterized in that: The calculating the similarity between the composite semantic vector and the image vector, and recommending matching materials based on the similarity ranking, includes: Calculate the cosine similarity between the composite semantic vector and the image vector, and recommend matching materials in descending order of the cosine similarity.

9. A news material recommendation device based on angle decoupling and cross-modal matching, characterized in that: include: Structured instruction building module, used to build structured instruction templates; A large model text parsing module is used to call the large model based on the structured instruction template to perform semantic parsing on the input text and extract multi-angle text elements; A feature encoding module, configured to perform cross-modal encoding on the multi-angle text elements and image materials to generate text vectors and image vectors; A weight optimization module, configured to generate a matching weight for each text element based on the text vector and the image vector through a dynamic weight optimization model; The matching recommendation module is used to generate a composite semantic vector according to the matching weight, calculate the similarity between the composite semantic vector and the image vector, and recommend matching materials based on the similarity ranking.

10. An electronic device, characterized in that: The electronic device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device implements the method according to any one of claims 1 to 8.