Content processing method and device, equipment and storage medium
By acquiring and analyzing images associated with text content, using machine learning models to determine the correlation between images and text, filtering out target images and correlating them with text content, solving the problem of lack of visual appeal in text content, and achieving more efficient information transmission and user experience.
Patent Information
- Application Number
- CN202510081035.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively solve the correlation problem between text content and images, resulting in the lack of visual appeal and intuitiveness of information when presenting text content.
By acquiring a plurality of candidate images associated with the text content, image description information is generated using the first model, and text content and image description information are provided to the second model to determine the correlation between the text content and the candidate image, thereby filtering out the target image and associating it to the text content.
It realizes the accurate filtering of target images with high correlation with text content from multiple candidate images and correlating them with text content, which improves the transmission rate of information and the expressiveness of text content, and enhances the user's visual experience.
Smart Images

Figure CN119988666A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to content processing methods, devices, apparatuses, and computer-readable storage media. Background Art
[0002] With the development of computer technology, various forms of electronic devices can greatly enrich people's daily life. For example, people can use electronic devices to process content. For example, a target image associated with text content can be selected from multiple images as a picture of the text content. How to process content is a focus of attention. Summary of the invention
[0003] In a first aspect of the present disclosure, a content method is provided. The method includes: obtaining multiple candidate images associated with text content; using a first model, generating image description information of the multiple candidate images; providing the text content and the image description information to a second model to determine the relevance between the text content and the multiple candidate images; based on the relevance between the text content and the multiple candidate images, determining a target image from the multiple candidate images; and associating the target image with the text content so that the target image is presented in association with the text content.
[0004] In a second aspect of the present disclosure, a device for content processing is provided. The device includes: an acquisition module configured to acquire multiple candidate images associated with text content; a generation module configured to generate image description information of multiple candidate images using a first model; a provision module configured to provide text content and image description information to a second model to determine the relevance between the text content and the multiple candidate images; a determination module configured to determine a target image from the multiple candidate images based on the relevance between the text content and the multiple candidate images; and an association module configured to associate the target image with the text content so that the target image is presented in association with the text content.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory, the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit. When the instructions are executed by the at least one processing unit, the device executes the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of the present disclosure, a computer program product is provided, which includes computer executable instructions, and when the instructions are executed by a processor, the method according to the first aspect of the present disclosure is implemented.
[0008] It should be understood that the contents described in this content section are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;
[0011] Figure 2A A flowchart showing an interactive process of content processing according to some embodiments of the present disclosure;
[0012] Figure 2B-2C An interface diagram according to some embodiments of the present disclosure is shown;
[0013] Figure 3 A schematic structural block diagram of an apparatus for content processing according to some embodiments of the present disclosure is shown;
[0014] Figure 4 A block diagram of an electronic device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0016] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this article, and any type of embodiment may be included under any section / subsection. In addition, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0017] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.
[0018] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects are subject to the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user knows and confirms. Accordingly, when implementing each embodiment of the present disclosure, the type, scope of use, usage scenario, etc. of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method can vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.
[0019] In this specification and the embodiments, if personal information processing is involved, it will be processed on the premise of having a legal basis (such as obtaining the consent of the subject of personal information, or it is necessary to perform a contract, etc.), and will only be processed within the scope of regulations or agreements. If a user refuses to process personal information other than the necessary information for basic functions, it will not affect the user's use of basic functions.
[0020] Traditionally, although text content can convey the main content, it lacks visual appeal and intuitiveness of information, and it is difficult to meet the needs of users and readers for more vivid and interactive content.
[0021] The embodiment of the present disclosure proposes a content processing scheme. According to the scheme, multiple candidate images associated with text content are obtained; image description information of the multiple candidate images is generated using a first model; the text content and the image description information are provided to a second model to determine the relevance between the text content and the multiple candidate images; based on the relevance between the text content and the multiple candidate images, a target image is determined from the multiple candidate images; and the target image is associated with the text content so that the target image is presented in association with the text content.
[0022] Based on this approach, the embodiments of the present disclosure can accurately screen out target images with high relevance to text content from multiple candidate images, and associate the screened target images with the text content, so that the target images can be displayed together when presenting the text content, thereby ensuring that users can obtain key information from a visual level, effectively improving the information transmission rate and the expressiveness of the text content, and improving the user's visual experience.
[0023] Example Environment
[0024] Figure 1 1 is a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 As shown, example environment 100 may include electronic device 110 .
[0025] In this example environment 100, the electronic device 110 may run an application 120 that supports interface interaction. The application 120 may be any appropriate type of application for interface interaction, examples of which may include, but are not limited to, information applications, news applications, social applications, or other appropriate applications. The user 140 may interact with the application 120 via the electronic device 110 and / or its attached devices.
[0026] exist Figure 1 In the environment 100, if the application 120 is in an active state, the electronic device 110 can present an interface 150 for supporting interface interaction through the application 120.
[0027] The server 130 may obtain multiple candidate images associated with the text content. Further, the server 130 may select a target image with high relevance to the text content from the multiple candidate images, and associate the target image with the text content, so that the target image may be presented on the interface 150 in association with the text content on the electronic device 110.
[0028] In some embodiments, the electronic device 110 communicates with the server 130 to provide services for the application 120. The electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio receiver, an e-book device, a game device, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.).
[0029] The server 130 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. The server 130 may include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, etc. The server 130 may provide background services for the application 120 that supports virtual scenes in the electronic device 110.
[0030] A communication connection may be established between the server 130 and the electronic device 110. The communication connection may be established in a wired manner or a wireless manner. The communication connection may include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and the embodiments of the present disclosure are not limited in this respect. In the embodiments of the present disclosure, the server 130 and the electronic device 110 may implement signaling interaction through the communication connection between the two.
[0031] It should be understood that the structure and function of the various elements in the environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.
[0032] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.
[0033] Example Process
[0034] Figure 2A 1 is a flowchart showing a process 200 of content processing according to some embodiments of the present disclosure. The process 200 may be implemented at the server 130. Figure 1 Process 200 is described.
[0035] At block 210 , the server 130 obtains a plurality of candidate images associated with text content.
[0036] In some embodiments, the text content may be associated with any appropriate type of content, such as information content (news content), novel content, article content, and the like.
[0037] In some embodiments, the text content may be obtained from one content source, or may be obtained by summarizing content obtained from multiple content sources. The content source is the source of obtaining or generating content, which may be any appropriate network platform, traditional media, etc.
[0038] As an example, the text content may be generated based on a group of related content in multiple content sources. The group of related content may be a group of texts associated with a predetermined event. For example, the server 130 may obtain news text 1 corresponding to the event of the opening of the Olympic Games from network platform 1, and obtain news text 2 corresponding to the event of the opening of the Olympic Games from network platform 2, and then summarize news text 1 and news text 2 to generate summarized text content.
[0039] In order to improve the effectiveness and intuitiveness of information transmission, the server 130 can obtain multiple candidate images associated with the text content, and use these multiple candidate images as candidate images to be presented in association with the text content, that is, one or more target images associated with the text content are screened out from the multiple candidate images and presented together with the text content to improve the user's reading experience. For example, the target image and the text content can be combined to obtain the target content that is finally presented to the user. The target image can be presented at a predetermined position in the target content, such as in front of the text content, that is, it can be presented in the target content as the header image of the target content.
[0040] In some embodiments, the server 130 may obtain multiple candidate images associated with text content from multiple content sources, and the text content is generated based on a group of associated content in the multiple content sources. This group of associated content may be a group of texts associated with a predetermined event. As an example, the server 130 may obtain news content 1 corresponding to the event of the opening of the Olympic Games from network platform 1, wherein news content 1 includes news text 1 and news image 1, and obtain news content 2 corresponding to the event of the opening of the Olympic Games from network platform 2, wherein news content 2 includes news text 2 and news image 2. The summarized news obtained by summarizing news text 1 and news text 2 is the text content, and the multiple candidate images associated with the text content include news image 1 and news image 2.
[0041] Since the number of images included in the image set associated with the text content may be large, in order to ensure the quality of the images presented in association with the text content, the multiple candidate images may be partial images screened from the image set associated with the text content. As an example, the images associated with the text content may be an image set consisting of images associated with a predetermined event acquired based on various content sources.
[0042] Specifically, the server 130 may determine the evaluation information corresponding to the image set associated with the text content. In some embodiments, the evaluation information may be determined by at least one of the following: image parameters of the candidate image, and screen content of the candidate image.
[0043] In some embodiments, the image parameters may include but are not limited to aspect ratio, resolution, aesthetics, etc. For different devices, the requirements corresponding to the aspect ratio are different. The more the aspect ratio is adapted to the device, the higher the corresponding evaluation information. The higher the resolution, the higher the corresponding evaluation information. The higher the aesthetics, the higher the corresponding evaluation information. In some embodiments, the server 130 may use a predetermined model to determine the aesthetics corresponding to the multiple candidate images. The predetermined model may be any appropriate machine learning model for aesthetic evaluation, which is used to determine the aesthetics of the image.
[0044] In some embodiments, the screen content may include but is not limited to words, tables, line graphs, etc. If the candidate image includes more words, the corresponding evaluation information is lower. If the candidate image includes a table or a line graph, the corresponding evaluation information is lower.
[0045] Furthermore, the server 130 may determine a plurality of candidate images whose evaluation information meets a preset condition from the image set. For example, the server 130 may determine images whose evaluation information in the image set is higher than a threshold as the plurality of candidate images.
[0046] In block 220 , the server 130 generates image description information for a plurality of candidate images using the first model.
[0047] In some embodiments, the first model may be any suitable machine learning model, such as a multimodal language model.
[0048] The image description information may indicate the content included in the image, specifically, the objects included in the image, the scene corresponding to the image, the action or activity corresponding to the object, the atmosphere corresponding to the image, and the like.
[0049] In some embodiments, the server 130 may input the multiple candidate images into the first model to obtain image description information of the multiple candidate images output by the first model.
[0050] Since it is difficult to understand mixed text and image content, in order to prevent the first model from excessively hallucinating the reference text, that is, excessively associating the relationship between the text and the image, in other embodiments, the server 130 may provide the first model with a first candidate image from among multiple candidate images to determine the initial description text corresponding to the first candidate image. The first candidate image may be any image from among the multiple candidate images. The initial description text may be initial description information to be subsequently polished or supplemented to obtain image description information.
[0051] Further, the server 130 may provide the first model with an initial description text, a first candidate image, and a reference text corresponding to the first candidate image to generate image description information corresponding to the first candidate image. As an example, the reference text may correspond to the same content source as the first candidate image, such as the first candidate image and the reference text corresponding to the first candidate image are an image and text in news B of network platform A. In some embodiments, the text content may be generated based on the reference text corresponding to the first candidate image. For example, if a plurality of candidate images include a first candidate image 1 and a first candidate image 2, and the first candidate image 1 corresponds to reference text 1, and the second candidate image 2 corresponds to reference text 2, then the text content may be generated based on reference text 1 and reference text 2. Specifically, the server 130 may summarize reference text 1 and reference text 2 to generate text content.
[0052] In some embodiments, if the reference text and the first candidate image correspond to the same content source, the server 130 may input the initial description text, the first candidate image, and the reference text corresponding to the first candidate image into the first model in a predetermined first format to obtain image description information corresponding to the first candidate image output by the first model. The predetermined first format may indicate the position information of the reference text and the first candidate image in the target content corresponding to this content source. For example, the reference text is the text of news B (target content) of network platform A, and the first candidate image is an image in news B of network platform A. The first format may indicate that the image is located at the center of the text.
[0053] As an example, the server 130 may input the initial description text, the first candidate image, and the reference text corresponding to the first candidate image into the first model in the following example format:
[0054] Reference text:
[0055] [Beginning of reference text]
[0056] {news prefix}
[0057] {image}
[0058] {news suffix}
[0059] [END OF REFERENCE TEXT]
[0060] Image Description
[0061] [Beginning of image]
[0062] {image_caption_str1}
[0063] [End of image]
[0064] In the above input example, the first candidate image is located in the middle of the reference text, that is, the first part of the text content in the reference text is located before the first candidate image, and the second part of the text content in the reference text is located after the first candidate image. {news prefix} can be used to insert the first part of the text content, {news suffix} is used to insert the second part of the text content, {image} is used to insert the first candidate image itself, and {image_caption_str1} is used to insert the initial description text corresponding to the first candidate image.
[0065] In some embodiments, in order to fix the format and / or length of the output corresponding to the first model at different stages, the server 130 can give the first model some samples based on the in-context learning method, so that the first model can learn what format and / or length of image description to output for different stages, wherein the different stages of the first model include a first stage of generating an initial description text and a second stage of generating image description information, that is, the first model is configured to generate an initial description text and / or image description information based on the sample information, wherein the sample information at least indicates the format and / or length of the information to be output.
[0066] Specifically, before generating the initial description text, the server 130 may provide the first model with a first example, and the first example may indicate a first sample image and the first sample image corresponds to a first description text, so that the first model can learn based on the first example that when subsequently generating the initial description text, the initial description text is generated and output based on the format and / or length of the first description text. The number of first sample images and first description texts can be set as required.
[0067] The server 130 may also provide a second example to the first model before generating the image description information. The second example may indicate the initial description text corresponding to the second sample image, the second sample image and the reference text corresponding to the second sample image, and the second description text corresponding to the second sample image, so that the first model can learn based on the second example that when subsequently generating the image description information, the image description information is generated and output based on the format and / or length of the second description text. The number of second sample images and second description texts can be set as required. It should be noted that the method for generating the initial description text corresponding to the second sample image can be the same as the method for generating the initial description text corresponding to the first candidate image, which will not be described in detail here.
[0068] At block 230 , the server 130 provides the text content and the image description information to the second model to determine the relevance of the text content and the plurality of candidate images.
[0069] In some embodiments, the second model may be any appropriate machine learning model, such as a language model.
[0070] In some embodiments, the server 130 may input the text content and the image description information into the second model in a predetermined second format to determine the relevance between the text content and the plurality of candidate images.
[0071] As an example, the server 130 may input the text content and image description information into the first model in the following example format:
[0072] Text content:
[0073] [Beginning of text content]
[0074] {text}
[0075] [End of text]
[0076] Image Description
[0077] [Beginning of image]
[0078] {image_caption_str2}
[0079] [End of image]
[0080] In the above input example, {text} can be used to insert text content, and {image_caption_str2} is used to insert image description information.
[0081] In some embodiments, in some embodiments, the second model is configured to determine the relevance based on a prompt word, and the prompt word indicates at least one constraint information used to determine the relevance. The at least one constraint may include but is not limited to whether the candidate image can summarize the core content of the text content, whether the candidate image can summarize the expression theme of the text content, whether the candidate image can reflect the predetermined event corresponding to the text content and the object involved in the predetermined event, whether the object included in the candidate image does not meet the predetermined condition, etc.
[0082] In some embodiments, the second model is further configured to output an explanation text corresponding to the relevance. Specifically, the prompt word provided to the second model may instruct the second model to output information of the explanation text corresponding to the relevance.
[0083] In some embodiments, the prompt word may also indicate an output format corresponding to the output of the second model.
[0084] In some embodiments, the prompt word may also indicate the target number of candidate images with high relevance rankings output by the second model, that is, the second model outputs some images among multiple candidate images as alternatives to the target image.
[0085] In some embodiments, the prompt word may also indicate that a candidate image that meets a predetermined image condition is not selected as a candidate for the target image. The predetermined image condition may be a predetermined occlusion in the image, such as an emoticon, mosaic, large font occlusion, etc., to ensure the clarity of the target image.
[0086] In block 240 , the server 130 determines a target image from the plurality of candidate images based on the relevance of the text content to the plurality of candidate images.
[0087] As an example, the server 130 may determine an image whose correlation is greater than a threshold value among multiple candidate images as a target image.
[0088] As another example, the server 130 may sort the multiple candidate images in descending order of relevance. Further, the server 130 may determine a predetermined number of images that are ranked first as target images.
[0089] At block 250 , the server 130 associates the target image with the text content so that the target image is presented in association with the text content.
[0090] In some embodiments, the text content may be associated with the information content, and the target image is provided as a picture of the information content. The information content may also be referred to as news content. For example, the target image may be presented as a header image of the news content on the electronic device 110. For another example, the target image may be recommended as the top of a news recommendation list on the electronic device 110.
[0091] Figure 2B and Figure 2C An interface diagram according to some embodiments of the present disclosure is shown, and the interface diagram is presented by an interface of the electronic device 120 .
[0092] by Figure 2B As an example, the electronic device 120 may present an interface 200B, which may be a chat interface between the user 140 and the virtual object. Relevant news recommendation content may be presented in the interface 200B. Specifically, the interface 200B includes an image 201 and text 203, an image 202, and text 204. Text 203 may be a title or summary content of summary news A associated with the image 201. Summary news A may be generated by summarizing news associated with event 1 obtained from multiple content sources. Text 204 may be a title or summary content of summary news B associated with the image 202. Summary news B may be generated by summarizing news associated with event 2 obtained from multiple content sources.
[0093] In some embodiments, the electronic device may present a Figure 2C As an example, the predetermined operation may be an operation of clicking on the image 202 or clicking on the text 204. Figure 2C As shown, text content 207 can be presented in interface 200C, wherein the text content 207 is the text content corresponding to the summary news, and the image 206 associated with the text content 207 is presented as a header image of the summary news B in interface 200C, wherein the text 205 can be the title of the summary news B.
[0094] Based on this approach, the embodiments of the present disclosure can accurately screen out target images with high relevance to text content from multiple candidate images and associate them with the text content, then associate the selected target images with the text content, and display them together when presenting the text content, effectively enhancing the information transmission rate and the expressiveness of the text content, and improving the user's visual experience.
[0095] Example devices and equipment
[0096] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 3 A schematic structural block diagram of an apparatus 300 for content processing according to some embodiments of the present disclosure is shown. The apparatus 300 may be implemented as or included in the electronic device 110 discussed above. Each module / component in the apparatus 300 may be implemented by hardware, software, firmware, or any combination thereof.
[0097] like Figure 3 As shown, the device 300 includes an acquisition module 310, configured to acquire multiple candidate images associated with text content; a generation module 320, configured to generate image description information of multiple candidate images using a first model; a providing module 330, configured to provide text content and image description information to a second model to determine the relevance between the text content and the multiple candidate images; a determination module 340, configured to determine a target image from the multiple candidate images based on the relevance between the text content and the multiple candidate images; and an association module 350, configured to associate the target image with the text content so that the target image is presented in association with the text content.
[0098] In some embodiments, the text content is associated with the information content, and the target image is provided as an illustration of the information content.
[0099] In some embodiments, the acquisition module 310 is further configured to: acquire multiple candidate images associated with text content from multiple content sources, where the text content is generated based on a group of associated content in the multiple content sources.
[0100] In some embodiments, the generation module 330 is further configured to: provide a first candidate image from multiple candidate images to the first model to determine an initial description text corresponding to the first candidate image; and provide the initial description text, the first candidate image and a reference text corresponding to the first candidate image to the first model to generate image description information corresponding to the first candidate image.
[0101] In some embodiments, the text content is generated based on a reference text corresponding to the first candidate image.
[0102] In some embodiments, the first model is configured to generate initial description text and / or image description information based on sample information, wherein the sample information at least indicates a format and / or length of information to be output.
[0103] In some embodiments, the second model is further configured to output an explanation text corresponding to the association.
[0104] In some embodiments, the second model is configured to determine the relevance based on a prompt word, the prompt word indicating at least one item of constraint information used to determine the relevance.
[0105] In some embodiments, the acquisition module 310 is further configured to: determine evaluation information corresponding to the image set associated with the text content; and determine a plurality of candidate images whose evaluation information meets a preset condition from the image set.
[0106] In some embodiments, the evaluation information is determined based on at least one of the following: an image parameter of the candidate image; and a picture content of the candidate image.
[0107] The units included in the device 300 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 300 can be implemented at least in part by one or more hardware logic components. As an example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0108] Figure 4 4 shows a block diagram of an electronic device 400 in which one or more embodiments of the present disclosure may be implemented. It should be understood that Figure 4 The electronic device 400 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. Figure 4 The electronic device 400 shown can be used to implement Figure 1 An electronic device 110 is shown.
[0109] like Figure 4As shown, the electronic device 400 is in the form of a general electronic device. The components of the electronic device 400 may include, but are not limited to, one or more processors or processing units 410, a memory 520, a storage device 430, one or more communication units 440, one or more input devices 450, and one or more output devices 460. The processing unit 410 may be an actual or virtual processor and is capable of performing various processes according to a program stored in the memory 420. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 400.
[0110] The electronic device 400 typically includes a plurality of computer storage media. Such media may be any accessible media accessible to the electronic device 400, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 420 may be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 430 may be a removable or non-removable medium, and may include a machine-readable medium, such as a flash drive, a disk, or any other medium, which may be capable of being used to store information and / or data (e.g., training data for training) and may be accessed within the electronic device 400.
[0111] The electronic device 400 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 4 As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. The memory 420 may include a computer program product 425 having one or more program modules that are configured to perform various methods or actions of various embodiments of the present disclosure.
[0112] The communication unit 440 implements communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented with a single computing cluster or multiple computing machines that can communicate through a communication connection. Therefore, the electronic device 400 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0113] The input device 450 may be one or more input devices, such as a mouse, a keyboard, a tracking ball, etc. The output device 460 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 400 may also communicate with one or more external devices (not shown) through the communication unit 440 as needed, such as a storage device, a display device, etc., communicate with one or more devices that allow a user to interact with the electronic device 400, or communicate with any device that allows the electronic device 400 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0114] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0115] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices, equipment, and computer program products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0116] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0117] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0118] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple implementations of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some implementations as replacements, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0119] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein.
Claims
1. A content processing method, comprising: Obtaining multiple candidate images associated with the text content; Generate image description information of the plurality of candidate images using the first model; Providing the text content and the image description information to a second model to determine the relevance of the text content and the plurality of candidate images; Determining a target image from the plurality of candidate images based on the association between the text content and the plurality of candidate images; as well as The target image is associated with the text content so that the target image is presented in association with the text content.
2. The method according to claim 1, wherein the text content is associated with information content, and the target image is provided as an illustration of the information content.
3. The method according to claim 1, wherein obtaining a plurality of candidate images associated with the text content comprises: The plurality of candidate images associated with the text content are obtained from a plurality of content sources, wherein the text content is generated based on a set of associated content in the plurality of content sources.
4. The method according to claim 1, wherein generating image description information of the plurality of candidate images using the first model comprises: Providing a first candidate image among the plurality of candidate images to the first model to determine an initial description text corresponding to the first candidate image; as well as The initial description text, the first candidate image, and a reference text corresponding to the first candidate image are provided to the first model to generate the image description information corresponding to the first candidate image. The method according to claim 4 , wherein the text content is generated based on the reference text corresponding to the first candidate image.
6. The method according to claim 4, wherein the first model is configured to generate the initial description text and / or the image description information based on sample information, wherein the sample information at least indicates the format and / or length of the information to be output. The method according to claim 1 , wherein the second model is further configured to output an explanation text corresponding to the association. 8 . The method according to claim 7 , wherein the second model is configured to determine the association based on a prompt word, the prompt word indicating at least one item of constraint information used to determine the association.
9. The method according to claim 1, wherein obtaining a plurality of candidate images associated with the text content comprises: Determining evaluation information corresponding to the image set associated with the text content; as well as The plurality of candidate images whose evaluation information meets a preset condition are determined from the image set.
10. The method according to claim 9, wherein the evaluation information is determined based on at least one of the following: Image parameters of the candidate image; The picture content of the candidate image.
11. A device for content processing, comprising: An acquisition module, configured to acquire a plurality of candidate images associated with the text content; A generating module, configured to generate image description information of the plurality of candidate images using the first model; A providing module, configured to provide the text content and the image description information to a second model to determine the relevance of the text content and the plurality of candidate images; a determination module configured to determine a target image from the plurality of candidate images based on the association between the text content and the plurality of candidate images; as well as The associating module is configured to associate the target image with the text content so that the target image is presented in association with the text content.
12. An electronic device comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.
13. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 10 when executed by a processor.
14. A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.
Citation Information
Cited By
Multimedia conference data processing method based on artificial intelligence model and related device
CN120897097A