Method for acquiring video data, method and apparatus for training deep learning model

By processing text data associated with video data, identifying target words, and acquiring related video data, the problems of high cost and low quality in video data acquisition are solved, enabling the acquisition of high-quality video data and effective training of deep learning models.

CN115098730BActive Publication Date: 2026-02-06BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210796905.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2026-02-06
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

The high cost and low quality of acquiring video data result in poor effectiveness of processing it.

Method used

By processing the first text data associated with the first type of video data, candidate words and word categories are obtained. Based on the word categories, target words are determined from the candidate words, and target video data associated with the first type of video data is obtained from the second type of video data.

Benefits of technology

This improved the quality and relevance of target video data, reduced the cost of acquiring video data, and ensured the effectiveness of deep learning model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115098730B_ABST
    Figure CN115098730B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for obtaining video data, a training method, device and medium of a deep learning model, and a product, relating to the technical field of artificial intelligence such as knowledge graph, natural language processing and deep learning. The method for obtaining video data comprises: processing first text data associated with first type video data to obtain a candidate word and a word category corresponding to the candidate word; determining a target word from the candidate word based on the word category; and obtaining target video data associated with the first type video data from second type video data based on the target word.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, such as knowledge graph, natural language processing, and deep learning, and more particularly, to a method for obtaining video data, a training method and device for deep learning model, an electronic device, a medium, and a program product. BACKGROUND

[0002] In some cases, it is necessary to utilize video data for related processing, for example, it is necessary to utilize video data as training samples to train a deep learning model. However, the way of obtaining video data is costly, and the quality of the video data is low, which leads to poor effect of utilizing video data for related processing. SUMMARY

[0003] The present disclosure provides a method for obtaining video data, a training method and device for deep learning model, an electronic device, a storage medium, and a program product.

[0004] According to an aspect of the present disclosure, a method for obtaining video data is provided, including: processing first text data associated with first type video data to obtain candidate words and word categories corresponding to the candidate words; determining target words from the candidate words based on the word categories; and obtaining target video data associated with the first type video data from second type video data based on the target words.

[0005] According to another aspect of the present disclosure, a training method for deep learning model is provided, including: obtaining sample video data; and training a deep learning model using the sample video data, wherein the sample video data is obtained according to the above method for obtaining video data.

[0006] According to another aspect of the present disclosure, a device for obtaining video data is provided, including: a processing module, a determining module, and an obtaining module. The processing module is configured to process first text data associated with first type video data to obtain candidate words and word categories corresponding to the candidate words. The determining module is configured to determine target words from the candidate words based on the word categories. The obtaining module is configured to obtain target video data associated with the first type video data from second type video data based on the target words.

[0007] According to another aspect of the present disclosure, a training device for deep learning model is provided, including: an obtaining module and a training module. The obtaining module is configured to obtain sample video data. The training module is configured to train a deep learning model using the sample video data, wherein the sample video data is obtained according to the above device for obtaining video data.

[0008] According to another aspect of the present disclosure, an electronic device is provided, including at least one processor and a memory connected to the at least one processor in communication. Wherein the memory stores instructions executable by the at least one processor, the instructions are executed by the at least one processor to enable the at least one processor to perform any one or more of the above-described methods of obtaining video data, the training method of a deep learning model.

[0009] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, the computer instructions being used to cause the computer to perform any one or more of the above-described methods of obtaining video data, the training method of a deep learning model.

[0010] According to another aspect of the present disclosure, a computer program product is provided, including computer program sequences / instructions stored on at least one of a readable storage medium and an electronic device, the computer program sequences / instructions being executed by a processor to implement the steps of any one or more of the above-described methods of obtaining video data, the training method of a deep learning model.

[0011] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0012] The accompanying drawings are used to better understand the present scheme, and do not constitute a limitation on the present disclosure. Among them:

[0013] Figure 1 The system architecture of obtaining video data and / or training a deep learning model according to an embodiment of the present disclosure is schematically shown;

[0014] Figure 2 The flowchart of the method of obtaining video data according to an embodiment of the present disclosure is schematically shown;

[0015] Figure 3 The principle diagram of the method of obtaining video data according to an embodiment of the present disclosure is schematically shown;

[0016] Figure 4A The schematic diagram of a candidate word list according to an embodiment of the present disclosure is schematically shown;

[0017] Figure 4B The schematic diagram of a candidate word list according to another embodiment of the present disclosure is schematically shown;

[0018] Figure 4C The schematic diagram of a tree structure according to an embodiment of the present disclosure is schematically shown;

[0019] Figure 5 a flowchart of a method for training a deep learning model is shown schematically according to an embodiment of the present disclosure;

[0020] Figure 6 a block diagram of an apparatus for obtaining video data is shown schematically according to an embodiment of the present disclosure;

[0021] Figure 7 a block diagram of an apparatus for training a deep learning model is shown schematically according to an embodiment of the present disclosure; and

[0022] Figure 8 is a block diagram of an electronic device for performing obtaining video data and / or training of a deep learning model according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, which are shown by way of illustration. Various details of the embodiments of the present disclosure are described herein in order to provide a thorough understanding of the embodiments of the present disclosure. However, persons of ordinary skill in the art will readily appreciate that the embodiments of the present disclosure can be practiced without a number of the details set forth herein without departing from the scope of the embodiments of the present disclosure. In other instances, well-known features are not described in detail to avoid obscuring the embodiments of the present disclosure. Accordingly, the embodiments of the present disclosure are not limited to the details of the descriptions.

[0024] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used herein, the term "including" and "comprising" and the like are meant to be inclusive in a manner that there are no other non-mentioned items. Described herein are only exemplary embodiments of the present disclosure, but the present disclosure is not limited thereto.

[0025] All terms used herein, including technical and scientific terms, have the meanings commonly understood by one of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having meanings that are consistent with the context of the specification, and should not be interpreted in an idealized or excessively formal manner.

[0026] In the case of using expressions similar to "at least one of A, B, and C, etc.", it should generally be interpreted to include any of them alone, any combination of two or more of them, and the like. For example, "a system having at least one of A, B, and C" should be interpreted to include a system having A alone, a system having B alone, a system having C alone, a system having A and B together, a system having A and C together, a system having B and C together, and / or a system having A, B, and C together, etc.

[0027] Figure 1 a system architecture for obtaining video data and / or training of a deep learning model is shown schematically according to an embodiment of the present disclosure. It should be noted that, Figure 1The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0028] like Figure 1 As shown, the system architecture 100 according to this embodiment may include clients 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between clients 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0029] Users can use clients 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on clients 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0030] Clients 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers. Clients 101, 102, and 103 in this embodiment of the disclosure can, for example, run applications.

[0031] Server 105 can be a server providing various services, such as a backend management server supporting websites browsed by users using clients 101, 102, and 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the clients. Alternatively, server 105 can also be a cloud server, meaning server 105 has cloud computing capabilities.

[0032] It should be noted that the method for acquiring video data provided in this embodiment can be executed by server 105. Accordingly, the device for acquiring video data provided in this embodiment can be located in server 105. Alternatively, the method for training a deep learning model provided in this embodiment can be executed by server 105. Accordingly, the device for training a deep learning model provided in this embodiment can be located in server 105.

[0033] In an example, the client 101, 102, 103 can send a video data obtaining instruction to the server 105, the video data obtaining instruction can include the first type of video data. The server 105 processes the first type of video data to obtain the target video data in response to the video data obtaining instruction. Alternatively, the first type of video data is stored in the server 105, and the server 105 processes the first type of video data to obtain the target video data based on the video data obtaining instruction after receiving the video data obtaining instruction from the client 101, 102, 103.

[0034] In another example, the client 101, 102, 103 can send a deep learning model training instruction to the server 105, the deep learning model training instruction can include sample video data. The server 105 trains the deep learning model using the sample video data in response to the deep learning model training instruction. Alternatively, the sample video data is stored in the server 105, and the server 105 trains the deep learning model using the sample video data based on the deep learning model training instruction after receiving the deep learning model training instruction from the client 101, 102, 103.

[0035] It should be understood that Figure 1 The number of clients, networks and servers in

[0036] The method for obtaining video data and the method for training deep learning model according to the exemplary embodiments of the present disclosure will be described below in conjunction with the system architecture of Figure 1 Figures 2-5 The method for obtaining video data and the method for training deep learning model according to the exemplary embodiments of the present disclosure can be executed by, for example, the server as shown in Figure 1 Figure 1 The server as shown in

[0037] Figure 2 A flowchart of the method for obtaining video data according to an embodiment of the present disclosure is schematically shown.

[0038] As shown in Figure 2 The method 200 for obtaining video data according to the embodiment of the present disclosure can include, for example, operation S210 to operation S230.

[0039] In operation S210, the first text data associated with the first type of video data is processed to obtain candidate words and word categories corresponding to the candidate words.

[0040] In operation S220, the target word is determined from the candidate words based on the word categories.​​

[0041] In operation S230, target video data associated with the first type of video data is obtained from the second type of video data based on the target word.

[0042] For example, the first type of video data includes video data in a first language, and the second type of video data includes video data in a second language. The first language is English, and the second language is Chinese. Alternatively, the first language is Chinese, and the second language is English. For ease of illustration, embodiments of the present disclosure are described by taking English as the first language and Chinese as the second language.

[0043] For the first type of video data, the first text data associated therewith is English text, and processing the first text data obtains a plurality of candidate words and a word category corresponding to each candidate word. The candidate words include Chinese words, and each candidate word is a character or a word.

[0044] After obtaining the plurality of candidate words and the word category of each candidate word, at least one target word can be selected from the plurality of candidate words based on the word category, and the word category of the target word meets a preset word category, for example. The selected target word is an important word in the plurality of candidate words, and the target word can represent the video content of the first type of video data.

[0045] After obtaining the target word, the target word can be matched with each second type of video data in the plurality of second type of video data respectively, and the second type of video data matched with the target word is taken as target video data, for example, Chinese video. The target video data is associated with the video content of the first type of video data, for example, the target video data is consistent or similar with the video theme of the first type of video data.

[0046] According to embodiments of the present disclosure, the number of the first type of video data is usually large and the quality is high. In order to obtain the second type of video data, the first type of video data can be taken as a reference, the target word is obtained by processing the first type of video data, and the target video data is searched from the second type of video data based on the target word. The video content and the video quality of the target video data are consistent with the video content and the video quality of the first type of video data. Thus, the target video data meeting the requirements can be obtained, and the cost of obtaining the video data is reduced.

[0047] In another example of the present disclosure, the first text data is obtained based on any one or more of title data, description information, subtitle information, and voice information of the first type of video data.

[0048] For example, title data of the first type of video data can be determined as the first text data.

[0049] Alternatively, description information of the first type of video data can be determined as the first text data. The description information, for example, includes a theme of the first type of video data, or includes summary information of content of the first type of video data.

[0050] Alternatively, caption information of the first type of video data is recognized to obtain the first text data. For example, the caption information can be recognized by using an optical character recognition (OCR) method, and the recognized caption information can be determined as the first text data.

[0051] Alternatively, speech information of the first type of video data is recognized to obtain the first text data. For example, the speech information can be recognized by using a speech recognition technology to obtain text, and the recognized text can be determined as the first text data.

[0052] Embodiments of the present disclosure can obtain the first text data in various ways, improve flexibility, and improve accuracy of the first text data.

[0053] In another example of the present disclosure, the text type of the first text data, for example, includes a first text type, and the first text type, for example, is an English type.

[0054] The video library, for example, includes a plurality of second type of video data. For each second type of video data, second text data associated with the second type of video data can be obtained, and a text type of the second text data, for example, is a second text type, and the second text type, for example, includes a Chinese type.

[0055] For example, the text type of the first text data can be converted from the first text type to the second text type to obtain converted first text data. For example, the first text data can be translated to obtain the converted first text data. Then, the converted first text data is processed to obtain the target word.

[0056] Next, based on the target word and the second text data, target video data is obtained from the second type video data, and the second text data corresponding to the target video data matches the target word. For example, at least one second type video data can be selected from the plurality of second type video data as the target video data. For example, for each second type video data, if the target word is included in the second text data corresponding to the second type video data, it means that the second text data matches the target word, and the second type video data can be selected as the target video data. Alternatively, if the second text data corresponding to the second type video data is similar to the target word, it means that the second text data matches the target word, and the second type video data can be selected as the target video data. For example, the second text data corresponds to a first vector (or part of the words in the second text data corresponds to the first vector), and the target word corresponds to a second vector. If the vector distance between the first vector and the second vector is less than a preset distance, it can be determined that the second text data is similar to the target word.

[0057] For example, the second text data associated with the second type video data can be obtained based on any one or more of the title data, the description information, the subtitle information, and the voice information of the second type video data.

[0058] For example, the title data of the second type video data is determined as the second text data. Alternatively, the description information of the second type video data is determined as the second text data. Alternatively, the subtitle information of the second type video data is recognized to obtain the second text data. Alternatively, the voice information of the second type video data is recognized to obtain the second text data. It can be understood that the manner of obtaining the second text data is similar to the manner of obtaining the first text data, which will not be described here.

[0059] According to an embodiment of the present disclosure, after obtaining the target word based on the first type video data, the target word and the second text data of the second type video data are matched to select the target video data from the plurality of second type video data, thereby improving the relevance of the target video data and the first type video data. In the case of high-quality data, the target video data obtained is also high-quality data with high probability.

[0060] Figure 3 The principle diagram of the method for obtaining video data according to an embodiment of the present disclosure is schematically shown.

[0061] As shown in Figure 3 For the first type video data 310, the first text data 311 corresponding to the first type video data 310 is, for example, English data. The text type of the first text data 311 is converted from English type to Chinese type, thereby obtaining the converted first text data 312.

[0062] After obtaining the converted first text data 312, the converted first text data 312 is processed by using a sequence labeling manner to obtain a candidate word list 313. The candidate word list 313 includes, for example, a plurality of candidate words and a word category corresponding to each candidate word. The sequence labeling manner is, for example, a natural language processing technology.

[0063] For the plurality of candidate words in the candidate word list 313, at least one candidate word is selected as a target word 314 based on the word category corresponding to each candidate word. The word category corresponding to the target word 314 is, for example, a preset word category.

[0064] The target word 314 is, for example, a keyword of the first type of video data 310, and the target word 314 represents the main content of the first type of video data 310. Therefore, the target video data can be obtained based on the target word 314.

[0065] For example, for a plurality of second type of video data 320, 330, 340 in a video library, the plurality of second type of video data 320, 330, 340 are, for example, Chinese videos. The second text data 321 corresponding to the second type of video data 320, the second text data 331 corresponding to the second type of video data 330, and the second text data 341 corresponding to the second type of video data 340 are, for example, Chinese data. The target word 314 is matched with the second text data 321, the target word 314 is matched with the second text data 331, and the target word 314 is matched with the second text data 341. If it is determined that the target word 314 matches the second text data 331, the second type of video data 330 corresponding to the second text data 331 can be determined as the target video data. In an example, the target word 314 matches the second text data 331, for example, indicating that the target word 314 is included in the second text data 331.

[0066] According to an embodiment of the present disclosure, the target video data obtained based on the target word has relevance with the first type of video data, and data of high quality target video data is obtained based on the first type of video data of high quality.

[0067] Figure 4A An illustrative diagram of a candidate word list according to an embodiment of the present disclosure is shown.

[0068] As Figure 4AAs shown, the candidate word list 413A includes at least a plurality of candidate words and a word category corresponding to each candidate word. The word category may include, for example, at least one of a first word category and a second word category. The candidate word list 413A may also include other information such as word length.

[0069] For example, sequence labeling can be used to process the transformed first text data to obtain multiple candidate words and the word category of each candidate word. Sequence labeling methods, for example, have functions such as word segmentation, word classification, and semantic understanding.

[0070] For example, performing word segmentation on the converted first text data "Step by step. Cut golden potatoes into classic mashed potatoes." yields multiple candidate words: "step by step", "step", "de", ".", "will", "golden", "potato", "cut into", "classic", "mashed potatoes", etc.

[0071] In one approach, each candidate word can be classified to obtain a first word category corresponding to each candidate word. For example, multiple categories can be pre-defined, and each candidate word can be predicted to belong to at least one of these categories. These pre-defined categories could include, for example, "people," "works," "objects," "organizations," "culture," "time," "food," "scenes / events," "lifestyle," "sensory characteristics," "quantifiers," "particles," "prepositions," "vocabulary terms," ​​"modifiers," and so on. A classification model can be used to predict the category of each candidate word to determine at least one first word category corresponding to each candidate word from the pre-defined multiple categories.

[0072] In another approach, semantic understanding can be performed on each candidate word to obtain a second word category corresponding to each candidate word. For example, knowledge graph technology can be used for semantic understanding to mine the second word category. Taking the candidate word "potato" as an example, semantic understanding of "potato" can be performed to obtain the superordinate concept or higher-level attribute of "potato". For example, the superordinate concept or higher-level attribute of "potato" can be obtained as the category "potato", and the superordinate concept or higher-level attribute of "potato" can be obtained as the category "food".

[0073] For example, the word category corresponding to the target word includes at least one of the following: noun category, scene category, and sensory feature category. For instance, the word category corresponding to the target word includes a first word category and a second word category. The first word category corresponding to the target word includes, for example, at least one of the following: a first noun category, a first scene category, and a first sensory feature category. The second word category corresponding to the target word includes, for example, at least one of the following: a second noun category, a second scene category, and a second sensory feature category.

[0074] In an example, the candidate word and the first word category of the candidate word can be obtained, and then the target word is determined from the candidate word based on the first word category. The first word category corresponding to the target word includes at least one of the following, for example: a first noun category, a first scene category, a first sensory feature category. The first noun category includes “food category”, for example. The first scene category represents a motion scene or a motion event of the first type of video data, for example. The first scene category indicates that the target word belongs to a verb, which can represent a motion scene or a motion event, for example, as shown in Figure 4A The first sensory feature category includes a color feature, a state feature, and the like, for example. The first sensory feature category indicates that the target word belongs to an adjective, which includes a color adjective, a state adjective, and the like, for example, as shown in Figure 4A The first sensory feature category can include “sensory features”, as shown.

[0075] In another way, the candidate word and the second word category of the candidate word can be obtained, and then the target word is determined from the candidate word based on the second word category. The second word category corresponding to the target word includes at least one of the following, for example: a second noun category, a second scene category, a second sensory feature category. The second noun category includes “food category”, “potato”, “mashed potatoes”, and the like, for example. The second scene category represents a domain scene, a motion scene or a motion event to which the first type of video data belongs, for example. The second scene category indicates that the target word belongs to a verb, which can represent a domain scene, a motion scene or a motion event, for example, as shown in Figure 4A The second scene category includes “life category” and “scene event”, for example. The “life category” represents a domain scene to which the first type of video data belongs, and the “scene event” represents a motion scene or a motion event of the first type of video data. The second sensory feature category includes a color feature, a state feature, and the like, for example. The second sensory feature category indicates that the target word belongs to an adjective, which includes a color adjective, a state adjective, and the like, for example, as shown in Figure 4A The second sensory feature category can include “sensory features”, as shown.

[0076] In some cases, the identification result of the first word category or the identification result of the second word category may be misidentified or missed, and therefore, the target word can be determined based on any one of the first word category and the second word category, that is, the first word category of the determined target word is a first preset category, or the second word category of the determined target word is a second preset category. The first preset category is a first noun category, a first scene category, or a first sensory feature category, for example. The second preset category is a second noun category, a second scene category, or a second sensory feature category, for example.

[0077] In one approach, the target word can be determined based on the first word category. For example, if the first word category of a candidate word is the first noun category, the first scene category, or the first sensory feature category, then that candidate word is taken as the target word. Figure 4A As shown, the first word category for the candidate words "potato" and "mashed potatoes" is the first noun category "food," so "potato" and "mashed potatoes" can be used as target words. Similarly, the first word category for the candidate word "cut into" is the first scene category "scene event," so "cut into" can be used as a target word. Likewise, the first word category for the candidate word "golden" is the first sensory feature category "sensory feature," so "golden" can be used as a target word.

[0078] In another approach, the target word can be determined based on the second word category. For example, when the second word category of a candidate word is a second noun category, a second scene category, or a second sensory feature category, the candidate word is identified as the target word. Figure 4A As shown, the second word categories for the candidate words "potato" and "mashed potatoes" include the second noun category "food," "potato," and "mashed potatoes," making them suitable as target words. Similarly, the second word categories for the candidate words "golden" and "cut into" include the second scene category "lifestyle" and "scene event," making them suitable as target words. Of course, the second word category for the candidate word "golden" can also include the second sensory feature category "sensory feature," making it a suitable target word. It is understood that the second word category for a target word can include multiple categories; for example, the second word category for the target word "golden" could include the second scene category "lifestyle" and the second sensory feature category "sensory feature."

[0079] In another approach, M1 target words can be determined from multiple candidate words based on a first word category. If the number of M1 target words is small or does not meet the actual needs, N1 target words can be determined from multiple candidate words based on a second word category. Then, after removing duplicate words from the M1 and N1 target words, the remaining words are taken as the final target words. M1 and N1 are, for example, integers greater than 0.

[0080] In some cases, the target word can be determined based on the first word category. For example, the first word category of the target word is the first preset category. The first preset category is, for example, the first noun category, the first scene category, or the first sensory feature category.

[0081] For example, when the first word category of the candidate word is the first noun category, the first scene category, or the first sensory feature category, the candidate word is determined as the target word. For example, M2 target words are determined from the plurality of candidate words based on the first word category, N2 target words are determined from the plurality of candidate words based on the second word category, and the target words that overlap in the M2 target words and the N2 target words are determined as the final target words. M2 and N2 are, for example, integers greater than 0. For example, if the M2 target words include “potato” and “potato mash”, and if the N2 target words include “potato” and “golden”, the overlapping word “potato” is determined as the target word. The process of determining M2 target words from the plurality of candidate words based on the first word category and determining N2 target words from the plurality of candidate words based on the second word category is described above and will not be repeated here.

[0082] In some cases, to reduce the amount of calculation, the first target word can be determined from the candidate words based on the first word category. When the number of the first target words is less than the preset number, the second target word is determined from the remaining candidate words based on the second word category. The remaining candidate words are the words in the candidate words other than the first target words.

[0083] For example, the first text data is subjected to word segmentation processing to obtain a plurality of candidate words. Each candidate word is classified to obtain a first word category corresponding to each candidate word. Then, the first target word is determined from the plurality of candidate words based on the first word category. The first word category of the first target word is, for example, the first noun category, the first scene category, or the first sensory feature category.

[0084] If the number of the first target words is greater than or equal to the preset number, the first target words are taken as the final target words. If the number of the first target words is less than the preset number, for the remaining candidate words other than the first target words in the plurality of candidate words, semantic understanding can be performed on each of the remaining candidate words to obtain a second word category corresponding to each of the remaining candidate words, and then based on the second word category, a second target word is determined from the remaining candidate words, and the second word category of the second target word is, for example, a second noun category, a second scene category, or a second sensory feature category. Finally, the first target words and the second target words are determined as the final target words.

[0085] It can be understood that the first word category of the candidate words is first determined, the first target words are determined based on the first word category, and when the number of the first target words is small, the second word category of the remaining candidate words is then determined, and the second target words are determined based on the second word category, thereby reducing the amount of calculation consumed in the process of determining the target words.

[0086] According to another example of the present disclosure, candidate words with a third word category can be deleted from the plurality of candidate words, and the remaining candidate words are determined as target words. The third word category includes, for example, at least one of the following: a quantity word category, an auxiliary word category, a preposition category, a modifier category, and an abstract category. The third word category contains less semantic information and is generally difficult to represent the content of the first type of video data, so after deleting it, the remaining candidate words are more likely to represent the content of the first type of video data, and it is more appropriate to take them as target words.

[0087] In the above manner, a plurality of target words can be obtained, which include, for example, "golden", "potato", "cut into", "mashed potatoes", and the like. The plurality of target words can be matched with the second text data of the second type of video data, and when the second text data of the second type of video data includes one or more target words, the second type of video data can be taken as target video data.

[0088] Figure 4B An illustrative diagram of a candidate word list according to another embodiment of the present disclosure is schematically shown.

[0089] As Figure 4B shown, the candidate word list 413B includes at least a plurality of candidate words and a word category corresponding to each candidate word, and the word category includes, for example, at least one of a first word category and a second word category. The candidate word list 413B also includes, for example, word length and other content.

[0090] For example, the converted first text data "As spring comes, more and more people choose to go on a trip." is segmented to obtain multiple candidate words "As", "spring comes", ",", "more and more", "of", "people", "choose", "to", "go on a trip", ".".

[0091] Similar to the above, the target word is determined from the multiple candidate words based on at least one of the first word category and the second word category. The determined target word includes, for example, "spring comes", "people", "choose", and "go on a trip". Among them, the "time category" and the "task category" in the first word category can be the first noun category, and the "life category", "time stage", "role category", and "task" in the second word category can be the second noun category.

[0092] Figure 4C An illustrative diagram of a tree structure according to an embodiment of the present disclosure is shown.

[0093] As shown in Figure 4C , the tree structure 450 includes, for example, P nodes corresponding to P categories, P being an integer greater than 1, wherein each node corresponds to a category.

[0094] Taking the candidate word "potato" as an example, the semantic understanding of the candidate word is performed to obtain the standard word "potato" corresponding to the candidate word "potato". For example, a search is performed in the knowledge base to obtain the standard word "potato" corresponding to the candidate word "potato". In other cases, the standard word is the same as the candidate word, for example.

[0095] Next, a target branch structure 451 associated with the standard word is determined from the tree structure 450, for example, the target branch structure 451 includes the standard word "potato". The target branch structure 451 includes, for example, Q nodes corresponding to Q categories, Q being an integer less than or equal to P. The Q categories include, for example, "potato", "root and stem category", "vegetable category", and "food category". At least one of the Q categories is determined as the second word category, for example, "food category" and "potato" are determined as the second word category for the candidate word "potato".

[0096] After obtaining the target video data, the target video data can also be filtered based on the length of the video, the quality of the text, the content of the subtitles, etc. to obtain filtered high-quality target video data.

[0097] Video data, as a carrier, includes multi-modal information, which includes picture information in the video, text information (subtitle information, description information), audio information, and the like. Video data plays an important role in modern life. With the development of deep learning, the understanding of video content by deep learning models plays an increasingly important role, and large-scale pre-training technology has become a hot spot. Using video data with multi-modal information to train deep learning models has become an important task. Compared with image data and text data, video data (including visual information and text information) contains more rich dynamic information and can express more diverse content. Therefore, how to obtain video data for training deep learning models under the premise of low cost has become a problem to be solved.

[0098] In some cases, the existing training data set includes a large amount of English video data, but lacks Chinese video training samples.

[0099] In some cases, the link of the Chinese video can be obtained through network technology, and the video data is further cleaned by using rules, so as to obtain small-scale Chinese video data. However, this acquisition method is low in efficiency, and needs cumbersome rules for cleaning, resulting in uneven distribution of the acquired Chinese video data in categories, which affects the training accuracy of the deep learning model.

[0100] In order to obtain Chinese video data with uniform category distribution and high quality as training samples, the embodiments of the present disclosure take English video data (first category video data) with uniform category distribution as the basis, and obtain Chinese video data (target video data) through the above method. The English video data (first category video data) with uniform category distribution, for example, represents multi-category data including life category, sports category, news category, and the like in the English video data, and the data of each category is distributed relatively uniformly.

[0101] After obtaining the Chinese video data (target video data), it is used as sample video data to train the deep learning model, and the specific process is as shown in Figure 5

[0102] Figure 5 A flowchart of a training method of a deep learning model according to an embodiment of the present disclosure is schematically shown.

[0103] As shown in Figure 5 , the training method 500 of the deep learning model of the embodiments of the present disclosure may, for example, include operation S510 to operation S520.

[0104] In operation S510, sample video data is obtained.

[0105] In operation S520, the sample video data is used to train the deep learning model. ​

[0106] Exemplarily, the sample video data comprises the target video data as described above.

[0107] It can be understood that, since the quality of the first type of video data is good and the data categories are uniformly distributed, the category distribution of the target words obtained by processing the first type of video data is uniform, and the categories of the sample video data searched from the Chinese video library based on the target words are also relatively uniform.

[0108] Exemplarily, the deep learning model comprises a visual question answering model, a visual retrieval model, a generative model, etc.

[0109] For the visual question answering model, the question data and the video data are input into the model, the model obtains the answer data from the video data, and automatically outputs the answer data.

[0110] For the visual retrieval model, the text data is input into the model, and the model automatically matches the video data associated with the text data. Alternatively, the video data is input into the model, and the model automatically matches other similar video data.

[0111] For the generative model, the text data is input into the model, and the model automatically generates the video data. Alternatively, the video data is input into the model, and the model automatically generates the text data.

[0112] The method of the embodiments of the present disclosure mines the sample video data, expands the scale of the sample video data, ensures the quality and quantity of the sample video data, and enhances the fine-grained discrimination ability of the deep learning model.

[0113] Figure 6 A block diagram of an apparatus for obtaining video data according to an embodiment of the present disclosure is shown.

[0114] As shown in Figure 6 the apparatus 600 for obtaining video data according to the embodiments of the present disclosure comprises a processing module 610, a determination module 620 and an obtaining module 630.

[0115] The processing module 610 can be configured to process the first text data associated with the first type of video data to obtain candidate words and word categories corresponding to the candidate words. According to the embodiments of the present disclosure, the processing module 610 can perform the operation S210 described above with reference to Figure 2 and will not be described here again.

[0116] The determination module 620 can be configured to determine a target word from the candidate words based on the word categories. According to the embodiments of the present disclosure, the determination module 620 can perform the operation S220 described above with reference to Figure 2 and will not be described here again.

[0117] The obtaining module 630 can be configured to obtain target video data associated with the first type of video data from the second type of video data based on the target word. According to an embodiment of the present disclosure, the obtaining module 630 can perform the operation S230 described above, for example. Figure 2 The operation S230 described above will not be repeated here.

[0118] According to an embodiment of the present disclosure, the processing module 610 includes a conversion sub-module and a processing sub-module. The conversion sub-module is configured to convert the text type of the first text data from a first text type to a second text type to obtain converted first text data. The processing sub-module is configured to process the converted first text data in a sequence labeling manner to obtain a candidate word and a word category corresponding to the candidate word.

[0119] According to an embodiment of the present disclosure, the obtaining module 630 includes a first obtaining sub-module and a second obtaining sub-module. The first obtaining sub-module is configured to obtain second text data associated with the second type of video data, where the text type of the second text data is a second text type. The second obtaining sub-module is configured to obtain target video data from the second type of video data based on the target word and the second text data, where the second text data corresponding to the target video data matches the target word.

[0120] According to an embodiment of the present disclosure, the word category includes a first word category; and the processing module 630 includes a first word segmentation sub-module and a classification sub-module. The first word segmentation sub-module is configured to perform word segmentation processing on the first text data to obtain a candidate word. The classification sub-module is configured to classify the candidate word to obtain a first word category corresponding to the candidate word.

[0121] According to an embodiment of the present disclosure, the word category includes a second word category; and the processing module 630 includes a second word segmentation sub-module and a semantic understanding sub-module. The second word segmentation sub-module is configured to perform word segmentation processing on the first text data to obtain a candidate word. The semantic understanding sub-module is configured to perform semantic understanding on the candidate word to obtain a second word category corresponding to the candidate word.

[0122] According to an embodiment of the present disclosure, the semantic understanding sub-module includes a semantic understanding unit, a first determination unit, and a second determination unit. The semantic understanding unit is configured to perform semantic understanding on the candidate word to obtain a standard word corresponding to the candidate word. The first determination unit is configured to determine a target branch structure associated with the standard word from a tree structure, where the tree structure includes P nodes, the P nodes correspond to P categories, the target branch structure includes Q nodes, the Q nodes correspond to Q categories, P is an integer greater than 1, and Q is an integer less than or equal to P. The second determination unit is configured to determine at least one category in the Q categories as the second word category.

[0123] According to an embodiment of the present disclosure, the determining module 620 comprises a first determining sub-module, a second determining sub-module, and a third determining sub-module. The first determining sub-module is configured to determine, based on the first word category, the first target word from the candidate words, wherein the candidate words comprise the first target word and remaining candidate words; the second determining sub-module is configured to, in response to determining that the number of the first target words is less than a preset number, determine, based on the second word category of the remaining candidate words, a second target word from the remaining candidate words; and the third determining sub-module is configured to determine the first target word and the second target word as the target words.

[0124] According to an embodiment of the present disclosure, the word category corresponding to the target word comprises at least one of a noun category, a scene category, and a sensory feature category.

[0125] According to an embodiment of the present disclosure, the determining module 620 is further configured to delete, from the candidate words, a candidate word with a third word category, and determine the remaining candidate words as the target words, wherein the third word category comprises at least one of a quantity word category, an auxiliary word category, a preposition category, and a modifier category.

[0126] According to an embodiment of the present disclosure, the first text data is obtained in at least one of the following manners: determining title data of the first type of video data as the first text data; determining description information of the first type of video data as the first text data; recognizing subtitle information of the first type of video data to obtain the first text data; and recognizing speech information of the first type of video data to obtain the first text data.

[0127] According to an embodiment of the present disclosure, the first obtaining sub-module comprises at least one of a first determining unit, a second determining unit, a first recognizing unit, and a second recognizing unit. The first determining unit is configured to determine title data of the second type of video data as the second text data; the second determining unit is configured to determine description information of the second type of video data as the second text data; the first recognizing unit is configured to recognize subtitle information of the second type of video data to obtain the second text data; and the second recognizing unit is configured to recognize speech information of the second type of video data to obtain the second text data.

[0128] Figure 7 A block diagram of a training apparatus of a deep learning model according to an embodiment of the present disclosure is shown schematically.

[0129] As shown in Figure 7 The training apparatus 700 of the deep learning model according to an embodiment of the present disclosure comprises, for example, an obtaining module 710 and a training module 720.

[0130] The obtaining module 710 can be configured to obtain sample video data. According to an embodiment of the present disclosure, the obtaining module 710 can perform, for example, the operations described above with reference toFigure 5 The operation S510 described above will not be repeated here.

[0131] The training module 720 can be configured to train the deep learning model by using the sample video data. According to an embodiment of the present disclosure, the training module 720 may, for example, perform the operation S510 described above with reference to FIG. 5. Figure 5 The operation S520 described above will not be repeated here.

[0132] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure and application of the user's personal information comply with the relevant laws and regulations, necessary security measures are taken, and the public order and good customs are not violated.

[0133] In the technical solution of the present disclosure, the authorization or consent of the user is obtained before the user's personal information is acquired or collected.

[0134] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0135] According to an embodiment of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, the computer instructions being used to make a computer execute any one or more of the methods for acquiring video data and the methods for training a deep learning model described above.

[0136] According to an embodiment of the present disclosure, a computer program product is provided, including computer programs / instructions, which, when executed by a processor, implement any one or more of the methods for acquiring video data and the methods for training a deep learning model described above.

[0137] Figure 8 is a block diagram of an electronic device for executing the method for acquiring video data and / or the method for training a deep learning model according to an embodiment of the present disclosure.

[0138] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device 800 is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0139] As Figure 8As shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0140] A plurality of components in the device 800 are connected to the I / O interface 805, including an input unit 806 such as a keyboard, a mouse, and the like; an output unit 807 such as various types of displays, speakers, and the like; the storage unit 808 such as a magnetic disk, an optical disk, and the like; and a communication unit 809 such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0141] The computing unit 801 can be various general and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The computing unit 801 performs various methods and processes described above, such as any one or more of the methods of obtaining video data, the methods of training a deep learning model. For example, in some embodiments, any one or more of the methods of obtaining video data, the methods of training a deep learning model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of any one or more of the methods of obtaining video data, the methods of training a deep learning model described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform any one or more of the methods of obtaining video data, the methods of training a deep learning model by any other appropriate means, such as by means of firmware.

[0142] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0143] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of any one or more of a general purpose computer, special purpose computer, or other programmable apparatus to produce a machine, such that the program code, when executed by the processor or controller, implements the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0144] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0145] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and any of a variety of input devices, including keyboards, mice, trackballs, microphones, touch screens, touch pads, etc., by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0146] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0147] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server can arise by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0148] It should be understood that various forms of flow shown above can be used, with steps reordered, added, or removed. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without limitation, as long as the desired results of the technology disclosed in the present disclosure are achieved.

[0149] The specific embodiments described above are not intended to limit the scope of the present disclosure. Those skilled in the art will understand that various modifications, combinations, sub-combinations, and alternatives can be made to the specific embodiments without departing from the spirit and principles of the present disclosure. Any further modifications, changes, improvements, and the like that come within the spirit and scope of the present disclosure should be considered as falling within the scope of the present disclosure.

Claims

1. A method for acquiring video data, comprising: Process the first text data associated with the first type of video data to obtain candidate words and word categories corresponding to the candidate words; Based on the word category, the target word is determined from the candidate words; as well as Based on the target words, obtain target video data associated with the first type of video data from the second type of video data; The word categories include a first word category and a second word category; The process of processing the first text data associated with the first type of video data to obtain candidate words and word categories corresponding to the candidate words includes: performing word segmentation on the first text data to obtain the candidate words; classifying the candidate words to obtain the first word category corresponding to the candidate words; performing semantic understanding on the candidate words to obtain the second word category corresponding to the candidate words; converting the text type of the first text data from a first text type to a second text type to obtain converted first text data; and processing the converted first text data using sequence labeling to obtain the candidate words and word categories corresponding to the candidate words. The step of obtaining target video data associated with the first type of video data from the second type of video data based on the target word includes: obtaining second text data associated with the second type of video data, wherein the text type of the second text data is the second text type; and obtaining the target video data from the second type of video data based on the target word and the second text data, wherein the second text data corresponding to the target video data matches the target word, and the second text data corresponding to the target video data contains the target word or is similar to the target word; The target word category conforms to the preset word category.

2. The method according to claim 1, wherein, The step of performing semantic understanding on the candidate words to obtain the second word category corresponding to the candidate words includes: Semantic understanding is performed on the candidate words to obtain the standard words corresponding to the candidate words; Determine a target branch structure associated with the standard words from a tree structure, wherein the tree structure includes P nodes, each corresponding to one of P categories, and the target branch structure includes Q nodes, each corresponding to one of Q categories, where P is an integer greater than 1 and Q is an integer less than or equal to P; and At least one of the Q categories is determined as the second word category.

3. The method according to claim 1, wherein, The step of determining the target word from the candidate words based on the word category includes: Based on the first word category, a first target word is determined from the candidate words, wherein the candidate words include the first target word and the remaining candidate words; In response to determining that the number of the first target words is less than a preset number, a second target word is determined from the remaining candidate words based on the second word category of the remaining candidate words; and The first target word and the second target word are determined as the target word.

4. The method according to claim 1, wherein, The word category corresponding to the target word includes at least one of the following: noun category, scene category, and sensory feature category.

5. The method according to any one of claims 1-4, wherein, The step of determining the target word from the candidate words based on the word category includes: Remove candidate words belonging to the third word category from the candidate words, and determine the remaining candidate words as the target words. The third word category includes at least one of the following: quantifier category, auxiliary word category, preposition category, and modifier category.

6. A method for training a deep learning model, comprising: Acquire sample video data; as well as Using the sample video data, a deep learning model is trained. The sample video data is obtained by the method according to any one of claims 1-5.

7. An apparatus for acquiring video data, comprising: The processing module is used to process the first text data associated with the first type of video data to obtain candidate words and word categories corresponding to the candidate words; The determining module is used to determine the target word from the candidate words based on the word category; as well as The acquisition module is used to acquire target video data associated with the first type of video data from the second type of video data based on the target words; The word categories include a first word category and a second word category; The processing module includes: a second word segmentation submodule, used to segment the first text data to obtain the candidate words; and a semantic understanding submodule, used to perform semantic understanding on the candidate words to obtain the second word category corresponding to the candidate words; The processing module includes: a first word segmentation submodule, used to segment the first text data to obtain the candidate words; and a classification submodule, used to classify the candidate words to obtain the first word category corresponding to the candidate words; The processing module includes: a conversion submodule, used to convert the text type of the first text data from a first text type to a second text type to obtain the converted first text data; and a processing submodule, used to process the converted first text data using sequence labeling to obtain the candidate words and the word categories corresponding to the candidate words. The acquisition module includes: a first acquisition submodule, configured to acquire second text data associated with the second type of video data, wherein the text type of the second text data is the second text type; and a second acquisition submodule, configured to acquire target video data from the second type of video data based on the target words and the second text data, wherein the second text data corresponding to the target video data matches the target words, and the second text data corresponding to the target video data contains the target words or is similar to the target words; The target word category conforms to the preset word category.

8. The apparatus according to claim 7, wherein, The semantic understanding submodule includes: A semantic understanding unit is used to perform semantic understanding on the candidate words to obtain standard words corresponding to the candidate words; A first determining unit is configured to determine a target branch structure associated with the standard words from a tree structure, wherein the tree structure includes P nodes, the P nodes corresponding to P categories, the target branch structure includes Q nodes, the Q nodes corresponding to Q categories, P is an integer greater than 1, and Q is an integer less than or equal to P; and The second determining unit is used to determine at least one of the Q categories as the second word category.

9. The apparatus according to claim 7, wherein, The determining module includes: The first determining submodule is configured to determine a first target word from the candidate words based on the first word category, wherein the candidate words include the first target word and the remaining candidate words; The second determining submodule is configured to, in response to determining that the number of the first target words is less than a preset number, determine the second target word from the remaining candidate words based on the second word category of the remaining candidate words; and The third determining submodule is used to determine the first target word and the second target word as the target word.

10. The apparatus according to claim 7, wherein, The word category corresponding to the target word includes at least one of the following: noun category, scene category, and sensory feature category.

11. The apparatus according to any one of claims 7-10, wherein, The determining module is also used for: Remove candidate words belonging to the third word category from the candidate words, and determine the remaining candidate words as the target words. The third word category includes at least one of the following: quantifier category, auxiliary word category, preposition category, and modifier category.

12. An apparatus for training a deep learning model, comprising: The acquisition module is used to acquire sample video data; as well as The training module is used to train a deep learning model using the sample video data. The sample video data is obtained by the apparatus according to any one of claims 7-11.

13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.

15. A computer program product comprising a computer program / instructions, characterized in that, The computer program / instructions are stored on at least one of a readable storage medium and an electronic device, and when executed by a processor, the computer program / instructions implement the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Information recommendation method and device, server and storage medium

    CN110334283A

  • Information processing method and device, electronic equipment, storage medium and program product

    CN114491149A