Media resource generation method and device, equipment and storage medium
By acquiring text materials and media resources of the target object, generating text description fragments and editing them, the problem of a single media resource generation method is solved, diversified media resource generation is realized, and the user experience is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2026-03-24
AI Technical Summary
Existing technologies use a single method for generating media resources, which fails to meet the diverse needs of users.
By acquiring the text and media resources of the target object, text description fragments are generated, and these fragments are used to edit and process the media resources to generate the target media resources.
It has enriched the ways in which media resources are generated, improved generation efficiency and user experience, and met users' needs for diverse media resources.
Smart Images

Figure CN121722929A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of data processing, and particularly relates to a media resource generation method and device, equipment and a storage medium. BACKGROUND
[0002] With the continuous development of media resource processing technology, people's functional requirements related to media resources are increasingly diversified. At present, the media resource generation mode for target objects is relatively single, and cannot meet the user's demand for media resources.
[0003] Therefore, how to enrich the generation mode of media resources to meet the user's demand for media resources is a technical problem to be solved at present. SUMMARY
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a media resource generation method, device, equipment and storage medium, which can enrich the generation mode of media resources.
[0005] In a first aspect, the embodiments of the present disclosure provide a media resource generation method, and the method comprises:
[0006] Obtaining text material and media resource material of a target object; wherein the text material comprises attribute description text of the target object, and the media resource material comprises at least one of picture material and video material;
[0007] Generating a text description segment of the target object based on the text material; wherein the text description segment is used to describe the characteristics of the target object;
[0008] Editing the media resource material by using the text description segment to generate a target media resource corresponding to the target object; wherein the target media resource is used to show the target object.
[0009] In an optional implementation, the generating a text description segment of the target object based on the text material comprises:
[0010] If the media resource material comprises at least one picture material, generating a text description segment corresponding to each of a plurality of preset dimensions and key recommendation information based on the text material; wherein the key recommendation information is extracted from the text description segment;
[0011] Correspondingly, the editing the media resource material by using the text description segment to generate a target media resource corresponding to the target object comprises:
[0012] The at least one picture material is edited by using the text description segment and the key recommendation information corresponding to each of the plurality of preset dimensions, and a graphic-text resource corresponding to the target object is generated.
[0013] In an alternative embodiment, before the at least one picture material is edited by using the text description segment and the key recommendation information corresponding to each of the plurality of preset dimensions, and a graphic-text resource corresponding to the target object is generated, the method further comprises:
[0014] determining a feature similarity between a first picture material in the at least one picture material and key recommendation information corresponding to a first preset dimension in the plurality of preset dimensions;
[0015] Correspondingly, the at least one picture material is edited by using the text description segment and the key recommendation information corresponding to each of the plurality of preset dimensions, and a graphic-text resource corresponding to the target object is generated, comprising:
[0016] if the feature similarity between the first picture material and the key recommendation information corresponding to the first preset dimension satisfies a preset similarity condition, the key recommendation information corresponding to the first preset dimension is synthesized into the first picture material to obtain a first edited picture material;
[0017] a graphic-text resource corresponding to the target object is generated based on the first edited picture material and a text description segment corresponding to the first preset dimension.
[0018] In an alternative embodiment, before the feature similarity between the first picture material in the at least one picture material and the key recommendation information corresponding to the first preset dimension in the plurality of preset dimensions is determined, the method further comprises:
[0019] an image feature vector of the first picture material in the at least one picture material is obtained, and a text feature vector of the key recommendation information corresponding to the first preset dimension in the plurality of preset dimensions is obtained;
[0020] Correspondingly, the feature similarity between the first picture material in the at least one picture material and the key recommendation information corresponding to the first preset dimension in the plurality of preset dimensions is determined, comprising:
[0021] based on the image feature vector of the first picture material and the text feature vector of the key recommendation information corresponding to the first preset dimension, the feature similarity between the first picture material and the key recommendation information corresponding to the first preset dimension is determined.
[0022] In an optional implementation, if the feature similarity between the first picture material and the key recommendation information corresponding to the first preset dimension satisfies a preset similarity condition, the key recommendation information corresponding to the first preset dimension is synthesized into the first picture material to obtain a first edited picture material, including:
[0023] If the feature similarity between the first picture material and the key recommendation information corresponding to the first preset dimension satisfies a preset similarity condition, the key recommendation information corresponding to the first preset dimension is synthesized into the first picture material based on a preset picture editing template to obtain a first edited picture material.
[0024] In an optional implementation, the generating of the target object corresponding graphic-text resource based on the first edited picture material and the text description segment corresponding to the first preset dimension includes:
[0025] The first graphic-text resource card is generated based on the first edited picture material and the text description segment corresponding to the first preset dimension, and the first graphic-text resource card is used to display the target object for the first preset dimension;
[0026] The target object corresponding graphic-text resource is generated based on the first graphic-text resource card.
[0027] In an optional implementation, the media resource material includes a video material, and the editing processing of the media resource material by using the text description segment to generate the target object corresponding target media resource includes:
[0028] The voice segment and / or the subtitle content segment are generated based on the text description segment, and the voice segment and / or the subtitle content segment are synthesized into the video material to generate the target object corresponding video resource.
[0029] In an optional implementation, after the editing processing of the media resource material by using the text description segment to generate the target object corresponding target media resource, the method further includes:
[0030] The target media resource is encapsulated into a message and sent to a preset message queue, and the preset message queue is used to transmit the target media resource to a resource display end.
[0031] In a second aspect, the disclosure provides a media resource generation device, and the device includes:
[0032] The acquisition module is configured to acquire text material and media resource material of a target object, wherein the text material comprises attribute description text of the target object, and the media resource material comprises at least one of picture material and video material.
[0033] The first generation module is configured to generate a text description segment of the target object based on the text material, wherein the text description segment is used to describe a feature of the target object.
[0034] The second generation module is configured to edit the media resource material by using the text description segment to generate a target media resource corresponding to the target object, wherein the target media resource is used to show the target object.
[0035] In a third aspect, the embodiments of the present disclosure further provide an electronic device, comprising a processor, a memory for storing executable instructions of the processor, and the processor is configured to read the executable instructions from the memory and execute the instructions to implement the media resource generation method provided by the embodiments of the present disclosure.
[0036] In a fourth aspect, the embodiments of the present disclosure further provide a computer readable storage medium, which stores a computer program for executing the media resource generation method provided by the embodiments of the present disclosure.
[0037] The technical solution provided by the embodiments of the present disclosure has the following advantages compared with the prior art.
[0038] In the media resource generation method provided by the embodiments of the present disclosure, first, text material and media resource material of a target object are acquired, wherein the text material comprises attribute description information of the target object, and the media resource material comprises at least one of picture material and video material; then, a text description segment of the target object is generated based on the text material, and the text description segment is used to describe a feature of the target object; and then, the media resource material is edited by using the text description segment to generate a target media resource corresponding to the target object. It can be seen that the embodiments of the present disclosure can edit the media resource material by using the text description segment corresponding to the text material of the target object to generate the target media resource corresponding to the target object, thereby enriching the generation mode of the media resource. BRIEF DESCRIPTION OF DRAWINGS
[0039] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.
[0040] Figure 1A flowchart of a media resource generation method provided by an embodiment of the present disclosure is shown in FIG. 1.
[0041] Figure 2 A schematic diagram of a first image-text resource card provided by an embodiment of the present disclosure is shown in FIG. 2.
[0042] Figure 3 A flowchart of another media resource generation method provided by an embodiment of the present disclosure is shown in FIG. 3.
[0043] Figure 4 A structural diagram of a media resource generation apparatus provided by an embodiment of the present disclosure is shown in FIG. 4.
[0044] Figure 5 A structural diagram of an electronic device provided by an embodiment of the present disclosure is shown in FIG. 5. DETAILED DESCRIPTION
[0045] Embodiments of the present disclosure will be described in more detail by referring to the drawings. Although certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms, and should not be construed as being limited to the embodiments set forth herein, but rather, these embodiments are provided so as to more completely and thoroughly understand the present disclosure. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only, and are not intended to limit the scope of protection of the present disclosure.
[0046] It is understood that each step recited in the method embodiments of the present disclosure can be executed in different orders, and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.
[0047] The term “comprising” and variations thereof as used herein are open-ended, that is, “including but not limited to”. The term “based on” is “based, at least in part, on”. The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments”. Related definitions are given throughout the description.
[0048] It is noted that the terms “first”, “second”, and the like in the present disclosure are merely used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0049] It is noted that the terms “one”, “multiple” in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise explicitly stated in the context, it should be understood as “one or more”.
[0050] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0051] Media resources refer to information content transmitted through various media formats, such as text, images, audio, video, and graphic media. Currently, the methods for generating media resources for specific target audiences are relatively limited and can no longer meet users' needs for media resources.
[0052] Therefore, this disclosure provides a media resource generation method. Specifically, firstly, text materials and media resource materials of a target object are obtained, wherein the text materials include attribute description information of the target object, and the media resource materials include at least one of image materials and video materials; then, a text description fragment of the target object is generated based on the text materials, the text description fragment being used to describe the characteristics of the target object; next, the media resource materials are edited using the text description fragment to generate target media resources corresponding to the target object. This disclosure can generate corresponding target media resources for a target object based on the text materials and media resource materials of the target object, for displaying the target object. Therefore, this disclosure enriches the methods for generating media resources.
[0053] Based on this, the present disclosure provides a method for generating media resources, which will be described below with reference to specific embodiments.
[0054] Figure 1 This is a flowchart illustrating a media resource generation method provided in an embodiment of the present disclosure. This method can be executed by a media resource generation device, which can be implemented using software and / or hardware, and is generally integrated into an electronic device. Figure 1 As shown, the method includes:
[0055] S101: Obtain the text materials and media resources of the target object.
[0056] The text material includes attribute description text of the target object, and the media resource material includes at least one of image material and video material.
[0057] The media resource generation method provided in this disclosure can be applied to a media resource generation server. The media resource server is used to generate corresponding target media resources for a target object and send the target media resources to a resource display terminal, which then displays the target media resources.
[0058] In this embodiment of the disclosure, media resources refer to information content that can be transmitted through various media forms, such as text resources, image resources, audio resources, video resources, and graphic resources.
[0059] The media resource generation method provided in this disclosure can be applied to various scenarios that require the production of media resources. For example, in the scenario of selling second-hand goods, corresponding target media resources are generated for second-hand goods to display them. In the scenario of recommending items, corresponding target media resources are generated for items to be recommended to display them. In the scenario of tourism, corresponding target media resources are generated for tourist attractions to promote them.
[0060] In this embodiment of the disclosure, the type of the target object may include item type, tourist attraction type, shop type, etc. For example, the target object may include items such as mobile phones, cars, and books in online stores, or tourist attraction A, etc. This embodiment of the disclosure does not limit the type of the target object.
[0061] In this embodiment of the disclosure, during the process of generating corresponding target media resources for a target object, the text material and media resource material of the target object are first obtained. The text material may include attribute description text of the target object. For example, assuming the target object is car A, the attribute description text corresponding to the target object may include car A's model information, configuration information, detection result information, etc.; assuming the target object is tourist attraction A, the attribute description text corresponding to the target object may include tourist attraction A's travel guide and other information.
[0062] The media resource materials may include at least one of image materials and video materials. For example, in the scenario of selling a used car, the media resource materials of the obtained car object may include image materials such as the exterior and interior of the vehicle, as well as video materials such as the engine test results video.
[0063] In one optional implementation, text materials and / or media resource materials of the target object can be obtained from the parameter information and detection report of the target object stored in the database. Specifically, the text materials can be obtained from the parameter information and detection report of the target object, and the media resource materials can be obtained from the image materials and video materials included in the detection report of the target object.
[0064] S102: Generate a text description fragment of the target object based on the text material.
[0065] The text description fragment is used to describe the characteristics of the target object.
[0066] In this embodiment of the disclosure, after obtaining the text material and media resource material of the target object, a text description fragment of the target object is generated based on the text material. The text description fragment can be used to describe the feature information of the target object. For example, assuming the target object is car A, the text description fragment corresponding to the target object may include feature information such as the usage, appearance, interior, power, configuration, detection items, and reasons for recommendation of car A.
[0067] In one optional implementation, since the text material may contain technical or difficult-to-understand parameter information or detection result information, after obtaining the text material of the target object, the text material is first processed by text conversion. Specifically, the technical terms, numbering information, etc. in the text material can be converted into easy-to-understand structured description text. Then, the structured description text is input into a preset model. After processing by the preset model, a text description fragment of the target object is generated.
[0068] S103: The media resource material is edited using the text description fragment to generate the target media resource corresponding to the target object.
[0069] The target media resource is used to display the target object.
[0070] In this embodiment of the disclosure, after generating a text description fragment of the target object based on the text material, the media resource material is edited using the text description fragment to generate the target media resource corresponding to the target object. The target media resource may include at least one of the image / text resources and video resources corresponding to the target object.
[0071] In this embodiment of the disclosure, the target media resources generated for the target object can be used to display in different scenarios. For example, in the scenario of selling second-hand goods, media resources corresponding to the items to be sold are displayed to the user so that the user can learn about the specifications, usage, and other information of the items to be sold based on the media resources; in the scenario of recommending items, media resources corresponding to the items are displayed to the user so that the user can learn about the features or advantages of the recommended items based on the media resources; and in the scenario of tourism, media resources of a certain place are displayed to the user so that the user can learn about the relevant details of that place based on the media resources.
[0072] Since image materials and video materials are two different types of materials, different editing methods can be used for image materials and video materials respectively. The specific processing methods will be introduced in subsequent embodiments and will not be elaborated here.
[0073] In practical applications, when displaying target media resources corresponding to a target object, different types of target media resources may have different generation times. For example, generating five different templates of text and image resources and video resources may take approximately 10 minutes in total. This means that while waiting for all the text and image resources and video resources to be generated before displaying them, users on the resource display end may experience a considerable amount of waiting time. Obviously, this method of synchronously generating target media resources is inefficient and provides a poor user experience on the resource display end.
[0074] Therefore, to improve the efficiency of media resource generation, target media resources can also be generated and displayed asynchronously. Specifically, after editing the media resource material using text description fragments to generate the target media resource corresponding to the target object, the target media resource is encapsulated into a message and sent to a preset message queue.
[0075] In this embodiment of the disclosure, after the media resource generation server generates the corresponding target media resource for the target object, it encapsulates the target media resource into a message and sends it to a preset message queue. Since the preset message queue has the characteristics of timely generation and timely transmission, that is, after the target media resource encapsulated into a message is sent to the preset message queue, the preset message queue will transmit the target media resource to the resource display end in a timely manner so that the resource display end can display the target object based on the target media resource.
[0076] For example, after generating five different image and text resources for the target object, the image and text resources are first encapsulated into messages and sent to a preset message queue. The preset message queue then transmits the image and text resources to the resource display terminal. While the image and text resources are displayed on the resource display terminal, video resources continue to be generated. Furthermore, after the video resources are generated, they are encapsulated into messages and sent to a preset message queue. The preset message queue then transmits the video resources to the resource display terminal, thus completing the generation and transmission of the target media resources asynchronously.
[0077] In the media resource generation method provided in this disclosure, firstly, text material and media resource material of a target object are obtained. The text material includes attribute description information of the target object, and the media resource material includes at least one of image material and video material. Then, a text description fragment of the target object is generated based on the text material. This text description fragment describes the characteristics of the target object. Next, the media resource material is edited using the text description fragment to generate the target media resource corresponding to the target object. Therefore, this disclosure embodiment can utilize the text description fragment corresponding to the text material of the target object to edit the media resource material to generate the target media resource corresponding to the target object, thereby enriching the methods for generating media resources.
[0078] In one optional implementation, when it is determined that the media resource material of the target object includes at least one image material, multiple text description fragments and key recommendation information corresponding to preset dimensions are first generated based on the text material.
[0079] Key recommendation information, also known as selling point information, can be displayed to users along with item information or images, allowing users to understand the advantages or features of the item based on the key recommendation information. In this embodiment, the key recommendation information can be extracted from text description fragments.
[0080] In this embodiment of the disclosure, multiple preset dimensions can be defined in advance for different types of target objects (or different media resource generation scenarios). When it is determined that the media resource material of the target object includes at least one image material, text description fragments and key recommendation information corresponding to the multiple preset dimensions are generated based on the text material.
[0081] For example, in the scenario of selling used cars, multiple preset dimensions such as usage, appearance, interior, power, configuration, inspection items, and reasons for recommendation can be predefined for the car object. If the obtained media resource material includes at least one image material, multiple text description fragments are generated based on the above multiple dimensions, and corresponding key recommendation information is extracted from each text description fragment.
[0082] In this embodiment of the disclosure, after generating text description fragments and key recommendation information corresponding to multiple preset dimensions based on text materials, at least one image material is edited using the text description fragments and key recommendation information corresponding to the multiple preset dimensions to generate image and text resources corresponding to the target object.
[0083] As can be seen, the embodiments of this disclosure generate multiple preset dimensions of text description fragments and key recommendation information for the image material of the target object, and use the multiple text description fragments and key recommendation information to edit and process the image material to generate the image and text resources corresponding to the target object, and then use the image and text resources to display the target object, thereby enriching the display method of the target object.
[0084] In one optional implementation, after editing at least one image material using text description fragments and key recommendation information, at least one image material can be matched with key recommendation information, and the successfully matched key recommendation information and image material can be combined to obtain the image and text resources corresponding to the target object.
[0085] Specifically, firstly, the feature similarity between the first image material in at least one image material and the key recommendation information corresponding to the first preset dimension in multiple preset dimensions is determined; if the feature similarity between the first image material and the key recommendation information corresponding to the first preset dimension meets the preset similarity condition, the key recommendation information corresponding to the first preset dimension is synthesized into the first image material to obtain the first edited image material; then, based on the first edited image material and the text description fragment corresponding to the first preset dimension, the image and text resources corresponding to the target object are generated.
[0086] In this embodiment of the disclosure, the first image material can be any one of at least one image material, and the first preset dimension can be any one of multiple preset dimensions. By determining the feature similarity between any image material and any key recommendation information, the key recommendation information whose feature similarity meets the preset similarity condition is synthesized into the image material to obtain the edited image material.
[0087] The preset similarity condition may include a feature similarity condition that is not less than a preset threshold. For example, when it is determined that the similarity between the first image material and the key recommendation information corresponding to the first preset dimension is greater than 70%, the key recommendation information corresponding to the first preset dimension is synthesized into the first image material to obtain the first edited image material.
[0088] In one optional implementation, before determining the feature similarity between the image material and the key recommendation information, the image feature vector of the first image material in at least one image material can be obtained by feature extraction, and the text feature vector of the key recommendation information corresponding to the first preset dimension in multiple preset dimensions can be obtained.
[0089] Among them, image feature vectors refer to a series of numerical values extracted from an image to describe the image's attribute information, while text feature vectors refer to numerical representations extracted from text data to describe the text's semantic and structural information.
[0090] In practical applications, a pre-trained neural network model can be used to extract features from the first image material to obtain its image feature vector, and to extract features from the key recommendation information to obtain its text feature vector. For example, in the pre-trained neural network model, the image feature vector of the first image material and the text feature vector of the key recommendation information can be mapped to the same vector space, allowing matching to be performed by calculating the similarity between the feature vectors.
[0091] In addition, after obtaining the image materials, they can be filtered. Specifically, at least one image material whose image clarity meets the preset clarity conditions can be identified as the image material to be edited, or at least one image material whose image aesthetics meet the preset aesthetic conditions can be identified as the image material to be edited. Furthermore, text description fragments and key recommendation information are used to edit the image materials that meet the above conditions to generate the image and text resources corresponding to the target object.
[0092] In this embodiment of the disclosure, after obtaining the image feature vector of the first image material and the text feature vector of the key recommendation information corresponding to the first preset dimension, the feature similarity between the first image material and the key recommendation information corresponding to the first preset dimension is determined based on the image feature vector of the first image material and the text feature vector of the key recommendation information corresponding to the first preset dimension.
[0093] Specifically, the feature similarity between the first image material and the key recommendation information corresponding to the first preset dimension can be determined by calculating the cosine distance between their image feature vectors and text feature vectors. Furthermore, when the feature similarity between the two satisfies a preset similarity condition, the key recommendation information corresponding to the first preset dimension is synthesized into the first image material to obtain the first edited image material.
[0094] In one optional implementation, when it is determined that the feature similarity between the first image material and the key recommendation information corresponding to the first preset dimension meets the preset similarity condition, the key recommendation information corresponding to the first preset dimension can also be synthesized into the first image material using a preset image editing template to obtain the first edited image material.
[0095] In practical applications, different preset image editing templates can be set for different business scenarios. These preset image editing templates can include information such as the display position, size, style, and format of key recommendation information in the image material. By using preset image editing templates, different types of graphic and text materials can be generated, thereby enriching the display functions of graphic and text materials, meeting users' browsing needs for different display formats of graphic and text materials, and improving the user experience.
[0096] In practical applications, after synthesizing the key recommendation information corresponding to the first preset dimension into the first image material to obtain the first edited image material, the target object can also be displayed using image and text resource cards to enrich the display methods of the target object.
[0097] Specifically, a first image-text resource card is generated based on the first edited image material and the corresponding text description fragment for the first preset dimension. This first image-text resource card is used to display the target object within the first preset dimension.
[0098] like Figure 2 The diagram shown is a schematic of a graphic resource card provided in an embodiment of this disclosure. Specifically, the resource card displays two graphic media resource cards. The first media resource card displays a first edited image material 201 and a text description fragment 202 corresponding to a first preset dimension. Specifically, the first edited image material 201 displays key recommendation information corresponding to the first preset dimension, such as "modern and dynamic," and also displays a brief description corresponding to the key recommendation information, such as "white body, streamlined design."
[0099] In one optional implementation, if the media resource material includes video material, then an audio segment and / or subtitle content segment are generated for the video material based on the text description segment, and the audio segment and / or subtitle content segment are synthesized into the video material to generate the video resource corresponding to the target object.
[0100] For example, assuming the acquired media resource material includes videos of the vehicle's exterior and interior, exterior description text and interior description text are generated based on text description fragments, serving as audio or subtitle content fragments of the video material. Furthermore, the audio fragments and / or subtitle content fragments are synthesized into the video material to obtain the vehicle's video resource. This allows users to more accurately or intuitively understand the content of the video material based on the subtitle content fragments or audio fragments in the video resource when it is subsequently displayed on the resource display terminal.
[0101] In practical applications, after generating text description fragments of the target object based on text materials, preset video editing templates can be used to generate audio fragments and / or subtitle content fragments based on the text description fragments. The preset video editing templates can be used to control the style or tone of the generated audio fragments.
[0102] In addition, after generating audio segments and / or subtitle content segments based on text description fragments, the duration information of video footage can be obtained, and the audio segments and / or subtitle content segments can be processed according to the duration information of video footage to ensure that the playback duration of the audio segments meets the playback requirements of the video footage (i.e., less than or equal to the playback duration of the video footage).
[0103] In one optional implementation, after generating multiple preset dimensions of text description fragments and key recommendation information based on text materials, or after generating audio fragments and / or subtitle content fragments based on text description fragments, data cleaning processing can be performed on the text description fragments, key recommendation information, audio fragments, subtitle content, etc., based on preset rules to obtain text data and language data that meet the synthesis requirements.
[0104] For example, text descriptions can be processed according to different word count or format requirements for different business scenarios to better suit the display needs of the current scenario. Additionally, since key recommendation information will be subsequently integrated into image assets, the word count of the key recommendation information can be controlled to ensure the composite image resources are aesthetically pleasing; for example, the word count of key recommendation information is generally limited to between 4 and 6 characters.
[0105] like Figure 3 The diagram shown is a flowchart of another media resource generation method provided in this embodiment of the present disclosure, specifically taking the example of media resource materials of the target object that simultaneously include image materials and video materials.
[0106] First, the text materials and media resources of the target object are retrieved from the database. The media resources include at least one image material and one video material. The database stores the parameter information, detection report, preset image editing templates, and other content of the target object.
[0107] Then, based on the text material, multiple preset dimensions of text description fragments and key recommendation information (also known as selling point information) are generated, where the key recommendation information can be extracted from the text description fragments. Additionally, audio fragments and / or subtitle content fragments are generated based on the text description fragments.
[0108] Next, the image feature vector of the first image material in at least one image material is obtained, and the text feature vector of the key recommendation information corresponding to the first preset dimension in multiple preset dimensions is obtained. Based on the image feature vector of the first image material and the text feature vector of the key recommendation information corresponding to the first preset dimension, the feature similarity between the first image material and the key recommendation information corresponding to the first preset dimension is determined.
[0109] If the feature similarity between the first image material and the key recommendation information corresponding to the first preset dimension meets the preset similarity condition, then the key recommendation information corresponding to the first preset dimension is synthesized into the first image material based on the preset image editing template to obtain the first edited image material, and the image and text resources corresponding to the target object are generated based on the first edited image material and the text description fragment corresponding to the first preset dimension.
[0110] After encapsulating the text and image resources corresponding to the target object into a message, the message is sent to a preset message queue. The preset message queue is then used to transmit the text and image resources to the resource display terminal. While displaying the text and image resources on the resource display terminal, audio clips and / or subtitle content clips are synthesized into the video material to generate the video resource corresponding to the target object.
[0111] After generating the video resource corresponding to the target object, the video resource corresponding to the target object is encapsulated into a message and sent to a preset message queue. The preset message queue is then used to transmit the video resource to the resource display terminal.
[0112] As can be seen, the embodiments of this disclosure can use the text description fragments corresponding to the text material of the target object to edit and process the media resource material to generate the target media resource corresponding to the target object, thereby enriching the way media resources are generated.
[0113] To implement the above embodiments, this disclosure also proposes a media resource generation apparatus. Figure 4 This is a schematic diagram of a media resource generation device provided in an embodiment of the present disclosure. The device can be implemented by software and / or hardware and is generally integrated into an electronic device. Figure 4 As shown, the device includes:
[0114] The acquisition module 401 is used to acquire text materials and media resource materials of the target object; wherein, the text materials include attribute description text of the target object, and the media resource materials include at least one of image materials and video materials;
[0115] The first generation module 402 is used to generate a text description fragment of the target object based on the text material; wherein the text description fragment is used to describe the characteristics of the target object;
[0116] The second generation module 403 is used to edit the media resource material using the text description fragment to generate the target media resource corresponding to the target object; wherein, the target media resource is used to display the target object.
[0117] In one optional implementation, the first generation module includes:
[0118] The first generation submodule is used to generate multiple text description fragments and key recommendation information corresponding to preset dimensions based on the text material if the media resource material includes at least one image material; wherein the key recommendation information is extracted from the text description fragments.
[0119] Accordingly, the second generation module includes:
[0120] The second generation submodule is used to edit the at least one image material using the text description fragments and key recommendation information corresponding to the multiple preset dimensions, and generate the image and text resources corresponding to the target object.
[0121] In one optional implementation, the second generation module further includes:
[0122] A determination submodule is used to determine the feature similarity between the first image material in the at least one image material and the key recommendation information corresponding to the first preset dimension in the plurality of preset dimensions;
[0123] Accordingly, the second generation submodule is specifically used for:
[0124] If the feature similarity between the first image material and the key recommendation information corresponding to the first preset dimension meets the preset similarity condition, then the key recommendation information corresponding to the first preset dimension is synthesized into the first image material to obtain the first edited image material;
[0125] The third generation submodule is used to generate the image and text resources corresponding to the target object based on the first edited image material and the text description fragment corresponding to the first preset dimension.
[0126] In one optional implementation, the second generation module further includes:
[0127] The acquisition submodule is used to acquire the image feature vector of the first image material in the at least one image material, and to acquire the text feature vector of the key recommendation information corresponding to the first preset dimension in the plurality of preset dimensions;
[0128] Accordingly, the determining submodule is specifically used for:
[0129] Based on the image feature vector of the first image material and the text feature vector of the key recommendation information corresponding to the first preset dimension, the feature similarity between the first image material and the key recommendation information corresponding to the first preset dimension is determined.
[0130] In one optional implementation, the second generation submodule is further configured to:
[0131] If the feature similarity between the first image material and the key recommendation information corresponding to the first preset dimension meets the preset similarity condition, then the key recommendation information corresponding to the first preset dimension is synthesized into the first image material based on the preset image editing template to obtain the first edited image material.
[0132] In one optional implementation, the third generation submodule is specifically used for:
[0133] Based on the first edited image material and the text description fragment corresponding to the first preset dimension, a first image and text resource card is generated; wherein, the first image and text resource card is used to display the target object according to the first preset dimension;
[0134] The graphic resources corresponding to the target object are generated based on the first graphic resource card.
[0135] In one optional implementation, the media resource material includes video material, and the second generation module includes:
[0136] The synthesis submodule is used to generate audio segments and / or subtitle content segments based on the text description segments, and to synthesize the audio segments and / or subtitle content segments into the video material to generate the video resource corresponding to the target object.
[0137] In one optional implementation, the method further includes:
[0138] The sending module is used to encapsulate the target media resource into a message and send it to a preset message queue; wherein, the preset message queue is used to transmit the target media resource to the resource display terminal.
[0139] The media resource generation apparatus provided in this disclosure acquires text material and media resource material of a target object. The text material includes attribute description information of the target object, and the media resource material includes at least one of image material and video material. A text description fragment of the target object is generated based on the text material, and this text description fragment is used to describe the characteristics of the target object. The media resource material is then edited using the text description fragment to generate a target media resource corresponding to the target object. Therefore, this disclosure embodiment can utilize the text description fragment corresponding to the text material of the target object to edit the media resource material to generate a target media resource corresponding to the target object, thereby enriching the methods for generating media resources.
[0140] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program / instructions, which, when executed by a processor, implements the media resource generation method in the above embodiments.
[0141] In addition, this disclosure also provides a media resource generation device, see [link to relevant documentation]. Figure 5 As shown, it may include:
[0142] The media resource generation device includes a processor 501, a memory 502, an input device 503, and an output device 504. The number of processors 501 in the device can be one or more. Figure 5Taking a processor as an example. In some embodiments of this disclosure, the processor 501, memory 502, input device 503, and output device 504 can be connected via a bus or other means, wherein, Figure 5 Taking the example of a connection between China and Israel via a bus.
[0143] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing of the media resource generation device by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc. In addition, the memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. The input device 503 can be used to receive input digital or character information, and to generate signal inputs related to user settings and function control of the media resource generation device.
[0144] Specifically in this embodiment, the processor 501 will load the executable files corresponding to the processes of one or more applications into the memory 502 according to the following instructions, and the processor 501 will run the applications stored in the memory 502 to realize the various functions of the media resource generation device.
[0145] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0146] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating media resources, characterized in that, include: Acquire text materials and media resource materials of the target object; wherein, the text materials include attribute description text of the target object, and the media resource materials include at least one of image materials and video materials; A text description fragment of the target object is generated based on the text material; wherein, the text description fragment is used to describe the characteristics of the target object; The media resource material is edited using the text description fragment to generate a target media resource corresponding to the target object; wherein, the target media resource is used to display the target object.
2. The method according to claim 1, characterized in that, The process of generating a text description fragment of the target object based on the text material includes: If the media resource material includes at least one image material, then multiple text description fragments and key recommendation information corresponding to preset dimensions are generated based on the text material; wherein, the key recommendation information is extracted from the text description fragments; Accordingly, the step of editing the media resource material using the text description fragment to generate the target media resource corresponding to the target object includes: The at least one image material is edited using text description fragments and key recommendation information corresponding to the multiple preset dimensions to generate image and text resources corresponding to the target object.
3. The method according to claim 2, characterized in that, Before editing and processing the at least one image material using the text description fragments and key recommendation information corresponding to the multiple preset dimensions to generate the image and text resource corresponding to the target object, the method further includes: Determine the feature similarity between the first image material in the at least one image material and the key recommendation information corresponding to the first preset dimension in the plurality of preset dimensions; Accordingly, the step of editing and processing the at least one image material using the text description fragments and key recommendation information corresponding to the multiple preset dimensions to generate the image and text resources corresponding to the target object includes: If the feature similarity between the first image material and the key recommendation information corresponding to the first preset dimension meets the preset similarity condition, then the key recommendation information corresponding to the first preset dimension is synthesized into the first image material to obtain the first edited image material; Based on the first edited image material and the text description fragment corresponding to the first preset dimension, the image and text resources corresponding to the target object are generated.
4. The method according to claim 3, characterized in that, Before determining the feature similarity between the first image material in the at least one image material and the key recommendation information corresponding to the first preset dimension in the plurality of preset dimensions, the method further includes: Obtain the image feature vector of the first image material in the at least one image material, and obtain the text feature vector of the key recommendation information corresponding to the first preset dimension in the plurality of preset dimensions; Accordingly, determining the feature similarity between the first image material in the at least one image material and the key recommendation information corresponding to the first preset dimension in the plurality of preset dimensions includes: Based on the image feature vector of the first image material and the text feature vector of the key recommendation information corresponding to the first preset dimension, the feature similarity between the first image material and the key recommendation information corresponding to the first preset dimension is determined.
5. The method according to claim 3, characterized in that, If the feature similarity between the first image material and the key recommendation information corresponding to the first preset dimension meets the preset similarity condition, then the key recommendation information corresponding to the first preset dimension is synthesized into the first image material to obtain the first edited image material, including: If the feature similarity between the first image material and the key recommendation information corresponding to the first preset dimension meets the preset similarity condition, then the key recommendation information corresponding to the first preset dimension is synthesized into the first image material based on the preset image editing template to obtain the first edited image material.
6. The method according to claim 3, characterized in that, The step of generating the image and text resources corresponding to the target object based on the first edited image material and the text description fragment corresponding to the first preset dimension includes: Based on the first edited image material and the text description fragment corresponding to the first preset dimension, a first image and text resource card is generated; wherein, the first image and text resource card is used to display the target object according to the first preset dimension; The graphic resources corresponding to the target object are generated based on the first graphic resource card.
7. The method according to claim 1, characterized in that, The media resource materials include video materials. The step of editing the media resource materials using the text description fragment to generate the target media resource corresponding to the target object includes: Based on the text description fragment, generate audio fragments and / or subtitle content fragments, and synthesize the audio fragments and / or subtitle content fragments into the video material to generate the video resource corresponding to the target object.
8. The method according to claim 1, characterized in that, After editing the media resource material using the text description fragment to generate the target media resource corresponding to the target object, the process includes: The target media resource is encapsulated into a message and sent to a preset message queue; wherein, the preset message queue is used to transmit the target media resource to the resource display terminal.
9. A media resource generation device, characterized in that, The device includes: The acquisition module is used to acquire text materials and media resource materials of the target object; wherein, the text materials include attribute description text of the target object, and the media resource materials include at least one of image materials and video materials; The first generation module is used to generate a text description fragment of the target object based on the text material; wherein the text description fragment is used to describe the characteristics of the target object; The second generation module is used to edit the media resource material using the text description fragment to generate the target media resource corresponding to the target object; wherein, the target media resource is used to display the target object.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a terminal device, cause the terminal device to perform the method as described in any one of claims 1-8.
11. A media resource generation device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as claimed in any one of claims 1-8.
12. A computer program product, characterized in that, The computer program product includes a computer program / instruction that, when executed by a processor, implements the method as described in any one of claims 1-8.