Object warehousing method and device, electronic equipment and medium

By automating the acquisition and matching of feature information from video data pages on websites, the problems of manpower consumption and timeliness in importing external video resources into the database have been solved, achieving efficient automatic import and improved user experience.

CN121880581APending Publication Date: 2026-04-17BEIJING IQIYI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING IQIYI TECH CO LTD
Filing Date
2025-12-31
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, the import of external video resources into the database requires manual review, which results in high human resource consumption and slow import timeliness, affecting user experience.

Method used

By extracting features from the data page of the object obtained from the first site, the first feature information is determined. Then, features are extracted from the playback source data of the object to be processed from the second site. Based on the feature information, the target data page is automatically matched and associated and stored in the database, replacing manual editing.

Benefits of technology

Automated warehousing has been achieved, reducing human resource consumption, improving warehousing efficiency and timeliness, and enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121880581A_ABST
    Figure CN121880581A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an object warehousing method and device, electronic equipment and a medium, and relates to the technical field of computers. The method comprises the steps that data pages of a plurality of objects are obtained from a first site, the data page of each object comprises detailed information of each object, feature extraction processing is carried out on each data page, and first feature information corresponding to each data page is determined; performing feature extraction processing on playing source data of a to-be-processed object from at least one second site, and determining second feature information of the to-be-processed object; determining a target data page matched with the to-be-processed object according to the first feature information and the second feature information; and associating the to-be-processed object with the target data page, and storing the to-be-processed object. According to the method, the data page of the object is automatically obtained to replace manual editing, and the first feature information and the second feature information are utilized to automatically associate the to-be-processed object to the target data page for automatic storage, so that the storage efficiency and timeliness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method, apparatus, electronic device, and medium for storing objects in a database. Background Technology

[0002] To provide a better user experience, some video websites, in addition to allowing users to search for their own video resources, often also support searching for resources from other websites (hereinafter referred to as external video resources). When searching for external video resources, methods such as web crawling or API calls from partners are generally used. To ensure video data quality and control risks, the acquired external video resources need to undergo manual review before being added to the database. For example, the external video resources are first submitted to the review backend for review by reviewers, and the approved playback sources are associated with a data page to complete the database entry. If there is no corresponding data page, the reviewers must create one. Manually creating data pages not only consumes considerable manpower, but also results in slower database entry times due to manpower limitations. This leads to some newly released movies and TV series from external websites not being added to the database in a timely manner, preventing users from finding relevant content promptly and thus affecting the user experience. Summary of the Invention

[0003] To solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, embodiments of the present invention provide an object storage method, apparatus, electronic device, and medium.

[0004] In a first aspect, embodiments of the present invention provide a method for storing objects in a database, including: Data pages of multiple objects are obtained from the first site. Each data page of an object includes detailed information about the object. Feature extraction processing is performed on each data page to determine the first feature information corresponding to each data page. Feature extraction processing is performed on the playback source data of the object to be processed from at least one second site to determine the second feature information of the object to be processed; Based on the first feature information and the second feature information, a target data page matching the object to be processed is determined; The object to be processed is associated with the target data page, and the playback source data of the object to be processed and the details on the target data page are written into the specified database.

[0005] Optionally, associate the object to be processed with the target data page, including: Bind the detailed information on the target data page as the description text of the object to be processed.

[0006] Optionally, after determining the first feature information corresponding to each of the data pages, the method further includes: establishing an index based on the first feature information corresponding to each of the data pages; The step of determining the target data page matching the object to be processed based on the first feature information and the second feature information includes: searching for target feature information matching the second feature information in the first feature information based on the index, and taking the data page corresponding to the target feature information as the target data page matching the object to be processed.

[0007] Optionally, the first feature information includes multiple first terms, and the second feature information includes multiple second terms, wherein the multiple second terms include the name of the object; Based on the index, the target feature information that matches the second feature information is searched in the first feature information, including: Determine the alias corresponding to the name of the object in the second feature information; Based on the index, search for candidate feature information that is the same as the name or alias of the object in the first feature information; For each first term in the candidate feature information, calculate the first similarity between the first term and the corresponding second term, and take the sum of the first similarities between each first term and the corresponding second term in the candidate feature information as the second similarity between the candidate feature information and the second feature information. Candidate feature information with a second similarity greater than or equal to the first threshold is taken as target feature information similar to the second feature information.

[0008] Optionally, after retrieving the data pages of multiple objects from the first site, the method further includes: The quality of the data pages of the multiple objects is evaluated to determine the first score of the data pages of each object. Based on the first score, the data pages of the multiple objects are filtered.

[0009] Optionally, the data page includes multiple detail fields and evaluation items. The multiple detail fields include at least one key field, and the evaluation items include one or more of the following: expected number of views, number of views already viewed, number of comments, and second rating. The number of detail fields is used to characterize the comprehensiveness of the data page's content, and the number of key fields is used to characterize the usefulness of the data page. The quality assessment of the data pages of the multiple objects, and the determination of the first score for the data pages of each object, includes: For each object's data page, determine the first target interval to which the number of detail fields included in the data page belongs from multiple first intervals, determine the second target interval to which the number of key fields belongs from multiple second intervals, and determine the third target interval to which the evaluation items belong from multiple third intervals; Calculate the weighted sum of the third score corresponding to the first target interval, the fourth score corresponding to the second target interval, and the fifth score corresponding to the third target interval, and use the weighted sum as the first score of the corresponding data page.

[0010] Optionally, before performing feature extraction processing on the playback source data of the object to be processed from at least one second site, the method further includes: obtaining playback source data of multiple candidate objects from at least one second site; and selecting the object to be processed from the multiple candidate objects according to a preset filtering rule.

[0011] Optionally, the filtering rules indicate one or more of the following: target type, target duration range, target release time, target theme, and completeness threshold of playback source data corresponding to the object to be processed; The step of selecting objects to be processed from the multiple candidate objects based on the playback source data of the multiple candidate objects and preset filtering rules includes: For each candidate object, based on the playback source data of the candidate object, determine one or more of the following: type, duration, release time, theme, and completeness of playback source data. From the plurality of candidate objects, select candidate objects that satisfy one or more of the following conditions, and use the candidate objects that satisfy one or more of the following conditions as objects to be processed: The type of the candidate object is consistent with the target type; The duration of the candidate is within the target duration range; The release time of the candidate object is consistent with the release time of the target object; The topic of the candidate object is consistent with the target topic; The completeness of the playback source data of the candidate object is greater than or equal to the completeness threshold.

[0012] Secondly, embodiments of the present invention provide an object storage device, comprising: The first extraction module is used to obtain data pages of multiple objects from the first site. Each data page of an object includes detailed information of the object. The module performs feature extraction processing on each data page to determine the first feature information corresponding to each data page. The second extraction module is used to perform feature extraction processing on the playback source data of the object to be processed from at least one second site, and to determine the second feature information of the object to be processed. The matching module is used to determine the target data page that matches the object to be processed based on the first feature information and the second feature information; The database entry module is used to associate the object to be processed with the target data page, and to write the playback source data of the object to be processed and the details on the target data page into the specified database.

[0013] Optionally, the inbound module is also used to bind the details information on the target data page as the descriptive text of the object to be processed.

[0014] Optionally, the device further includes a construction module, configured to: establish an index based on the first feature information corresponding to each of the data pages; The matching module is also used to: search for target feature information that matches the second feature information in the first feature information based on the index, and take the data page corresponding to the target feature information as the target data page for matching the object to be processed.

[0015] Optionally, the first feature information includes multiple first terms, and the second feature information includes multiple second terms, wherein the multiple second terms include the name of the object; The matching module is further configured to: determine the alias corresponding to the name of the object in the second feature information; search for candidate feature information that is the same as the name or alias of the object in the first feature information based on the index; calculate the first similarity between the first term and the corresponding second term for each first term in the candidate feature information, and take the sum of the first similarities between each first term and the corresponding second term in the candidate feature information as the second similarity between the candidate feature information and the second feature information; and take the candidate feature information with the second similarity greater than or equal to the first threshold as the target feature information similar to the second feature information.

[0016] Optionally, the device further includes a first filtering module, configured to: perform a quality assessment on the data pages of the plurality of objects, determine a first score for the data pages of each object, and filter the data pages of the plurality of objects based on the first score.

[0017] Optionally, the data page includes multiple detail fields and evaluation items. The multiple detail fields include at least one key field, and the evaluation items include one or more of the following: expected number of views, number of views already viewed, number of comments, and second rating. The number of detail fields is used to characterize the comprehensiveness of the data page's content, and the number of key fields is used to characterize the usefulness of the data page. The first filtering module is further configured to: determine the first target interval to which the number of detail fields included in the data page belongs from multiple first intervals, determine the second target interval to which the number of key fields belongs from multiple second intervals, and determine the third target interval to which the evaluation items belong from multiple third intervals; calculate the weighted sum of the third score corresponding to the first target interval, the fourth score corresponding to the second target interval, and the fifth score corresponding to the third target interval, and use the weighted sum as the first score of the corresponding data page.

[0018] Optionally, the device further includes a second filtering module, used to: obtain playback source data of multiple candidate objects from at least one second site; and filter out objects to be processed from the multiple candidate objects according to preset filtering rules.

[0019] Optionally, the filtering rules indicate one or more of the following: target type, target duration range, target publication time, target topic, and completeness threshold of playback source data corresponding to the object to be processed; the second filtering module is further used for: The step of selecting objects to be processed from the multiple candidate objects based on the playback source data of the multiple candidate objects and preset filtering rules includes: For each candidate object, based on the playback source data of the candidate object, determine one or more of the following: type, duration, release time, theme, and completeness of playback source data. From the plurality of candidate objects, select candidate objects that satisfy one or more of the following conditions, and use the candidate objects that satisfy one or more of the following conditions as objects to be processed: The type of the candidate object is consistent with the target type; The duration of the candidate is within the target duration range; The release time of the candidate object is consistent with the release time of the target object; The topic of the candidate object is consistent with the target topic; The completeness of the playback source data of the candidate object is greater than or equal to the completeness threshold.

[0020] Thirdly, embodiments of the present invention also provide an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to implement the object storage method provided in any embodiment of the present invention when executing the program stored in the memory.

[0021] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the object storage method provided in any embodiment of the present invention.

[0022] The technical solution provided in this embodiment of the method brings at least the following beneficial effects: The object entry method provided in this invention involves obtaining data pages of multiple objects from a first site, each data page including detailed information of the object; performing feature extraction processing on each data page to determine the first feature information corresponding to each data page; performing feature extraction processing on playback source data of the object to be processed from at least one second site to determine the second feature information of the object to be processed; determining the target data page matching the object to be processed based on the first and second feature information; associating the object to be processed with the target data page; and entering the object to be processed into the database. This method automates the acquisition of object data pages, replacing manual editing, reducing human resource consumption and costs. Furthermore, by utilizing the first and second feature information, this method automatically associates the object to be processed with the target data page for automatic entry into the database, improving entry efficiency and timeliness. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0024] Figure 1 A flowchart illustrating an embodiment of the object import method of the present invention is shown; Figure 2 A schematic diagram of a sub-process of the object entry method according to an embodiment of the present invention is shown; Figure 3 A flowchart illustrating another embodiment of the object storage method of the present invention is shown; Figure 4 A flowchart illustrating another embodiment of the object storage method of the present invention is shown; Figure 5 A schematic diagram of the object storage device according to an embodiment of the present invention is shown; Figure 6 A schematic diagram of the structure of an electronic device according to an embodiment of the present invention is shown. Detailed Implementation

[0025] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0026] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0027] Figure 1 A flowchart illustrating an embodiment of the object storage method of the present invention is shown. Figure 1 As shown, the method includes: Step S101: Obtain the data pages of multiple objects from the first site. Each object's data page includes the object's detailed information. Perform feature extraction processing on each data page to determine the first feature information corresponding to each data page.

[0028] In this context, a site refers to the address or location of a website, webpage, or network service on the Internet. In computer networks, a site can refer to all webpages or content under a domain name or IP address. A site can include various forms of content such as text, images, audio, and video. Users can access a site by entering its address or through a search engine. An object can include, but is not limited to, various forms of content such as text, images, audio, and video. For specific examples, an object can be an e-book, image, music, movie, or TV series. A data page is used to introduce an object, including collecting, organizing, and recording various detailed information. For example, in the case of an e-book, the data page includes the book title, author, table of contents, synopsis, publication date, and reviews. In the case of a video such as a movie or TV series, the data page includes the name, director, screenwriter, release date, synopsis, episode summaries, and cast list. The first site can be a site with comprehensive coverage and many data pages. Optionally, the first site for obtaining data pages may differ depending on the type of object selected. Optionally, multiple object data pages can be fetched from the first site using a custom crawler, or the first site can provide multiple object data pages by agreement with the first site.

[0029] Feature extraction is performed on each data page to extract the first feature information from the details displayed on the page. For example, by analyzing the structure of the data page, the first feature information can be extracted from the HTML structure. HTML (HyperText Markup Language) is a markup language that includes a series of tags, such as...<title>、< / title> . <title> The page title has been defined.< / title> This defines the actual content displayed on the webpage. This embodiment can be derived from tags. <title>、< / title> Extracting the first feature information. For example, the information page can be saved as an image, and the text on the page can be recognized using Optical Character Recognition (OCR) to convert it into a text file. The first feature information can then be extracted from this text file. For instance, in the case of an e-book, the first feature information extracted from the e-book's information page might include the book title, author, and publication date. In the case of a video, such as a movie or TV series, the first feature information extracted from the information page might include the title, director, release date, synopsis, and main actors.

[0030] In optional implementation scenarios, after obtaining the first feature information corresponding to the data pages of multiple objects, the first feature information corresponding to these multiple objects is stored in a database table for retrieval. As an optional example, as shown in Table 1 below, when the object is a video such as a movie or TV series, the information included in the first feature information, such as name, director, release time, plot summary, and main actors, is used as the column names of the database table.

[0031] Table 1:

[0032] Step S102: Perform feature extraction processing on the playback source data of the object to be processed from at least one second site to determine the second feature information of the object to be processed.

[0033] The second site can be any site other than the first site and its own site. The playback source data of the object to be processed refers to the relevant information displayed on the second site. For example, if the object to be processed is a video, the playback source data includes information such as the name, director, screenwriter, synopsis, and cast / crew list displayed on the second site. The content included in the playback source data of the object to be processed may be the same as or different from the content included in the information page obtained from the first site. Optionally, the playback source data of the object to be processed can be retrieved from the second site using a custom scraping program, or the second site can provide the playback source data of the object to be processed through an agreement.

[0034] Feature extraction processing is performed on the playback source data to extract second feature information. When the object to be processed is a video, the playback source data includes information such as the name, director, screenwriter, synopsis, and cast / crew list displayed on the second site. The name, director, screenwriter, and cast / crew list (which records the actors' names and the names of their roles) of the object to be processed are extracted from this playback source data as the second feature information. For example, an HTTP (Hypertext Transfer Protocol) request is sent to the second site to access it. The HTML structure of the webpage on the second site is obtained from the response to the HTTP request. The HTML structure is parsed, and features such as the `<html>` tag are extracted from the HTML tags. <title>、< / title> Extract the second feature information. Optionally, the fields included in the second feature information can be the same as or different from the fields included in the first feature information; no restrictions are imposed here. The fields included in the second feature information and the first feature information can be flexibly set according to the type of object (e.g., books, movies, music); no restrictions are imposed here either.

[0035] Step S103: Determine the target data page that matches the object to be processed based on the first feature information and the second feature information.

[0036] In this step, referring to Table 1, the first feature information includes fields such as object name, director, release date, plot summary, and main actors. The second feature information includes the name, director, screenwriter, and cast / crew table of the object to be processed. The same fields, such as name and director, included in the first and second feature information are matched sequentially to determine the target data page for the object to be processed.

[0037] Step S104: Associate the object to be processed with the target data page, and write the playback source data of the object to be processed and the details information on the target data page into the specified database.

[0038] In this step, the detailed information on the target data page is bound as the descriptive text of the object to be processed. For example, as shown in Table 2 below, the detailed information on the target data page (such as the names of the screenwriter, cast and crew, and their roles) is recorded in the corresponding entry for the object to be processed in Table 2. The playback source data of the object to be processed and the detailed information of the target data page are then written to a specified database. For example, Table 2 is written to a specified database.

[0039] Table 2:

[0040] The object entry method provided in this invention involves obtaining data pages of multiple objects from a first site, each data page including detailed information of the object; performing feature extraction processing on each data page to determine the first feature information corresponding to each data page; performing feature extraction processing on playback source data of the object to be processed from at least one second site to determine the second feature information of the object to be processed; determining the target data page matching the object to be processed based on the first and second feature information; associating the object to be processed with the target data page; and entering the object to be processed into the database. This method automates the acquisition of object data pages, replacing manual editing, reducing human resource consumption and costs. Furthermore, by utilizing the first and second feature information, this method automatically associates the object to be processed with the target data page for automatic entry into the database, improving entry efficiency and timeliness.

[0041] The object entry method provided in this invention can be applied to search or content recommendation scenarios. For example, based on this object entry method, at least one object to be processed from a second site is associated with the data page of a first site and written into a designated database. The objects in the designated database are not only numerous but also highly timely. During retrieval or recommendation, the designated database provides effective data support for search engines or content recommendation engines, improving the retrieval or content recommendation effect.

[0042] Taking video as an example, the object entry method provided in this embodiment of the invention constructs an automated entry process for external video. First, it automatically obtains the data page of the object to replace the data page manually created by the editor in the site. Then, for the entry of external video, it uses customized rules to identify the target data page corresponding to the video to be entered into the database, and automatically aggregates the video to be entered into the database with the target data page to complete the entire entry process, reducing labor costs and improving the efficiency and timeliness of video entry into the database.

[0043] In an optional embodiment, after obtaining the data pages of multiple objects from the first site in step S101, the data pages can be manually modified, such as adding, deleting, or updating information on the data pages. After the objects to be processed and the target data pages are entered into the database, the target data pages corresponding to the objects already entered into the database can be modified, or they can be deleted from the database.

[0044] In an optional embodiment, after determining the first feature information corresponding to each data page in step S101, the object storage method further includes: establishing an index based on the first feature information corresponding to each data page. An index is a data structure used to quickly locate the storage location of specific data in a database table. Indexes are typically stored using a B-tree or B+ tree structure (both B-trees and B+ trees are tree-like data structures). Each node contains a key value and a pointer to the next lower-level page. During retrieval, starting from the root node, the search value is compared with the key value, and the search proceeds layer by layer down the pointer until a matching leaf node is found, ultimately locating the data row. Optionally, the index can be established using SQL (Structured Query Language) or database management tools. For example, an index can be established using the standard SQL syntax CREATE INDEX (a function statement in SQL used to create an index in a database table). The basic format includes: CREATE INDEX index_name ON table_name(column_name1 [ASC|DESC], column_name2 [ASC|DESC], ...). The table_name is the name of the data table for which the index needs to be created, such as the name of table 1. Column names may include, for example, the object's name, director, release date, plot summary, and cast. After creating the index, based on this index, the target feature information matching the second feature information is searched in the first feature information, and the data page corresponding to the target feature information is used as the target data page for matching the object to be processed. For example, a query is initiated using an SQL query statement (such as SELECT * FROM tableWHERE column = 'value', which is used to retrieve records that meet specific conditions from a specified data table). Following the structure of a B-tree or B+ tree, the name in the "object name" column is searched for that is the same as the name of the object to be processed in the second feature information. A pointer is obtained to the data row containing the name that is the same as the name of the object to be processed in the second feature information. The data row is located and read using the pointer, and the data row is returned as the query result.

[0045] In an optional embodiment, the second feature information includes multiple terms, each term being the value of a specific field. For example, the field could be "director," and the value of that field (i.e., the term) could be "Zhang Moumou." Another example is the field could be the name of an object, and the value of that field (i.e., the term) could be, for instance, the title of a TV series.

[0046] Figure 2 This diagram illustrates a sub-process for object import according to an embodiment of the present invention. Figure 2 As shown, the process of searching for target feature information that matches the second feature information in the first feature information based on the index includes: Step S201: Determine the alias (e.g., the abbreviation of the TV series name) corresponding to the name of the object in the second feature information.

[0047] Step S202: Based on the index, search for candidate feature information in the first feature information that is the same as the name or alias of the object. For example, initiate a query using an SQL query statement (such as SELECT * FROM table WHERE column = 'value', which is used to retrieve records that meet specific conditions from a specified data table). Following the structure of a B-tree or B+ tree, search for a name in the "object name" column that is the same as the name or alias of the object to be processed in the second feature information. Obtain a pointer to the data row containing the name that is the same as the name of the object to be processed in the second feature information. Locate and read the data row using the pointer; the data in that data row is the candidate feature information.

[0048] Step S203: For each first term in the candidate feature information, calculate the first similarity between the first term and its corresponding second term. The sum of the first similarities between each first term and its corresponding second term in the candidate feature information is taken as the second similarity between the candidate feature information and the second feature information. For example, for a first term and its corresponding second term, convert the first term and the second term into vectors respectively, calculate the distance between the vector corresponding to the first term and the vector corresponding to the second term, and use this distance as the similarity between the first term and the second term.

[0049] Step S204: Select candidate feature information with a second similarity greater than or equal to the first threshold as target feature information similar to the second feature information.

[0050] In this embodiment, searching for candidate feature information that is the same as the name or alias of the object in the first feature information is called fuzzy retrieval, and determining the target feature information from the selected feature information is called precise retrieval. Fuzzy retrieval can narrow the scope of precise retrieval, improving matching efficiency and flexibility, while precise retrieval improves the accuracy of retrieval results.

[0051] Figure 3 A flowchart illustrating another embodiment of the object storage method of the present invention is shown. Figure 3 As shown, the method includes: Step S301: Obtain information pages for multiple objects from the first site. Each object's information page includes detailed information about the object. Evaluate the quality of the information pages for multiple objects and determine the first score for each object's information page.

[0052] Optionally, the information page includes multiple detail fields and evaluation items. Detail fields describe the basic information and details of the object. Taking a video as an example, detail fields include, but are not limited to: name, alias, director, screenwriter, release date, synopsis, episode summaries, cast and crew list, etc. Detail fields can be divided into general fields and key fields, with key fields being more important than general fields. As an optional example, key fields include name, director, release date, etc. Evaluation items represent online users' evaluations of the object. As an optional example, evaluation items include one or more of the following: expected number of views, number of views already viewed, number of comments, and second rating. The second rating is a score obtained by combining the evaluations of multiple online users on the first site. The number of detail fields represents the comprehensiveness of the information page's content, while the number of key fields represents the usefulness of the information page.

[0053] When evaluating the quality of each object's profile page, the page is assessed based on the number of detail fields, the number of key fields, and evaluation items to determine its first score. For example, it can be determined whether the page has too few detail fields, whether key fields are missing, and the page's expected view count, actual view count, number of comments, and second score to determine its corresponding first score. A higher first score indicates better page quality.

[0054] Optionally, multiple first intervals can be set, where the values ​​within each first interval represent the number of detail fields. Each first interval corresponds to a rating, which represents the rating corresponding to the number of detail fields belonging to the first interval. Multiple second intervals can be set, where the values ​​within each second interval represent the number of key fields. Each second interval corresponds to a rating, which represents the rating corresponding to the number of key fields belonging to the second interval. Multiple third intervals can be set, where a set of values ​​represents the expected number of views, the number of views, the number of comments, and the second rating, respectively. For example, a third interval could be [(a1, b1, c1, d1), (a2, b2, c2, d2)], where a1 and a2 represent the boundary values ​​for the expected number of views within the third interval, b1 and b2 represent the boundary values ​​for the number of views within the third interval, c1 and c2 represent the boundary values ​​for the number of comments within the third interval, and d1 and d2 represent the boundary values ​​for the second rating within the third interval. Each third interval corresponds to a rating, which represents the rating corresponding to the evaluation item belonging to the third interval. Then, determine the first target interval to which the number of detail fields included in the data page belongs from multiple first intervals, the second target interval to which the number of key fields belongs from multiple second intervals, and the third target interval to which the evaluation items belong from multiple third intervals; calculate the weighted sum of the third score corresponding to the first target interval, the fourth score corresponding to the second target interval, and the fifth score corresponding to the third target interval, and use the weighted sum as the first score of the corresponding data page. For example, determine the first score of the data page according to the following formula (1): P=w1*z1+w2*z2+w3*z3(1) Where P represents the first rating of the data page, w1, w2 and w3 represent the weights respectively, z1 represents the third rating, z2 represents the fourth rating and z3 represents the fifth rating.

[0055] Step S302: Filter the data pages of multiple objects based on the first score.

[0056] Optionally, in this step, data pages with a first rating less than a specified threshold can be deleted, and only data pages with a first rating greater than or equal to the specified threshold can be retained, thereby filtering out high-quality data pages.

[0057] Step S303: Perform feature extraction processing on each filtered data page to determine the first feature information corresponding to each data page. This step can be referenced... Figure 1 The embodiments shown are not described in detail here to avoid repetition.

[0058] Step S304: Obtain playback source data of multiple candidate objects from at least one second site, and select the object to be processed from the multiple candidate objects according to the preset filtering rules.

[0059] Optionally, the filtering rules indicate one or more of the following: target type, target duration range, target release time, target theme, and completeness threshold of playback source data for the object to be processed.

[0060] Taking videos as an example, video types include TV series, movies, animation, and variety shows. Video lengths include less than 30 minutes, more than 30 minutes but less than 1 hour, and more than 1 hour. The video's release date (e.g., its first broadcast date) is 2024 or 2025. The video's theme describes the subject matter, such as suspense or historical themes. The completeness of the playback source data characterizes the integrity of the playback source data. Completeness can be determined by the number of terms contained in the playback source data; the more terms, the higher the completeness. For example, if the completeness of the playback source data for video A is 90% and the completeness of the playback source data for video B is 80%, then the playback source data for video A is more complete than that for video B.

[0061] Taking books as an example, book types can be categorized by purpose or target audience. For instance, by purpose, book types include textbooks, reference books, and leisure reading materials; by target audience, book types include children's books, teen books, and adult books. Book themes can be divided by subject area, such as literature, lifestyle, education, and social sciences. The book's publication date (e.g., first publication date) is 2024 or 2025. The playback source data for the book is source data obtained from a second site (e.g., including book title, author, table of contents, synopsis, publication date, and reviews). The completeness of the playback source data can also be determined based on the number of terms contained within it.

[0062] From multiple candidate objects, select those that match one or more of the following criteria: target type, target duration range, target release time, target theme, and completeness threshold of playback source data. For example, from multiple candidate objects, select those that meet one or more of the following criteria as candidates to be processed: The type of the candidate object is consistent with the target type; The duration of the candidate is within the target duration range; The release time of the candidate is consistent with the target release time; The topic of the candidate matches the target topic; The completeness of the playback source data of the candidate object is greater than or equal to the completeness threshold.

[0063] As an optional example, taking videos as an example, if the filtering rules indicate that the target type of the object to be processed is TV series, the theme is suspense, the release time is 2025, and the completeness threshold is 80%, then among multiple candidate videos, TV series that were first broadcast in 2025, are suspenseful, and have a completeness of playback source data greater than or equal to 80% will be selected as objects to be processed.

[0064] Optionally, when the candidate is a TV series, shorter videos can be excluded based on the duration of a single episode, such as micro-dramas, spin-offs of the main series, and special versions. Optionally, filtering can also be performed based on the completeness of terms included in the playback source data, such as whether the cast and crew field is complete. In this embodiment of the invention, different filtering rules can be set for different sites, different types of objects, and application needs; this invention does not impose limitations. As an optional example, a second site's TV series channel library page contains a lot of micro-drama data. These relatively low-quality playback sources do not meet the definition of a main series and need to be filtered. Since the data source data crawled from this site includes the video duration, a threshold is set based on the video duration to filter out micro-drama objects.

[0065] Step S305: Perform feature extraction processing on the playback source data of the object to be processed from at least one second site to determine the second feature information of the object to be processed. This step can be referred to... Figure 1 The embodiments shown are not described in detail here to avoid repetition.

[0066] Step S306: Determine the target data page that matches the object to be processed based on the first feature information and the second feature information.

[0067] Playback source data from different secondary sites may contain content of the same object, so it is necessary to aggregate the playback source data of the same object to the corresponding target information page.

[0068] Step S307: Associate the object to be processed with the target data page, and write the playback source data of the object to be processed and the details of the target data page into the specified database. In this step, the details on the target data page are bound as the description text of the object to be processed (e.g., Table 2 above), and the details on the target data page are recorded in the entry corresponding to the object to be processed in Table 2. The playback source data of the object to be processed and the details of the target data page are then written into the specified database.

[0069] The object entry method provided in this invention involves obtaining data pages of multiple objects from a first site, each data page including detailed information of the object; performing feature extraction processing on each data page to determine the first feature information corresponding to each data page; performing feature extraction processing on playback source data of the object to be processed from at least one second site to determine the second feature information of the object to be processed; determining the target data page matching the object to be processed based on the first and second feature information; associating the object to be processed with the target data page; and entering the object to be processed into the database. This method automates the acquisition of object data pages, replacing manual editing, reducing human resource consumption and costs. Furthermore, by utilizing the first and second feature information, this method automatically associates the object to be processed with the target data page for automatic entry into the database, improving entry efficiency and timeliness.

[0070] Figure 4 A flowchart of an object import method according to another embodiment of the present invention is shown. Figure 4 In the illustrated embodiment, the objects imported into the database are movies and TV dramas. For example... Figure 4 As shown, the object entry method includes two data page feature retrieval services and production processes, and external site playback source aggregation processes.

[0071] The data page feature retrieval service and production process include: (1) Perform feature extraction processing on the data pages of each film and television drama in the entire network film and television database, obtain the first feature information corresponding to each data page, and save the obtained first feature information to the film and television feature database. Optionally, the data pages in the entire network film and television database can be obtained from the first site through a crawling program. The data pages are used to introduce the details of the film and television drama, such as including but not limited to the name of the film and television drama, director, screenwriter, release time, plot synopsis, episode plot, cast and crew list, expected number of views, number of views, number of comments, and second rating. The second rating is, for example, a score obtained by combining the evaluations of the object by multiple network users on the first site. Still optional, the first feature information includes: title (name of the film and television drama), alias, release time, channel (the channel is the entrance to the film and television drama category or theme, and different types or themes of film and television dramas are on different channels), region (e.g., the release region of the film and television drama), number of episodes, actors, director, etc.

[0072] (2) Based on the first feature information in the film and television feature library, establish a film and television data index.

[0073] The process of aggregating external playback sources includes: (1) Obtain the playback source data of the video to be processed, and preprocess the playback source data of the video to be processed, such as standardization processing. Standardization processing refers to processing the text into a unified and standardized form, such as deleting URL links and special characters in the playback source data.

[0074] (2) Feature extraction processing is performed on the playback source data of the preprocessed video to obtain the second feature information. (Reference) Figure 1 Step S102 of the illustrated embodiment will not be described again here.

[0075] (3) Based on the film and television data index, query candidate feature information similar to the second feature information from the film and television feature database. (Reference) Figure 2 Step S202 of the illustrated embodiment will not be described again here.

[0076] (4) Calculate the similarity between the candidate feature information and the second feature information. Candidate feature information with a similarity greater than or equal to the first threshold is considered as target feature information similar to the second feature information, and the data page corresponding to the target feature information is considered as the target data page. (Reference) Figure 2 Step S203 of the illustrated embodiment will not be described again here.

[0077] (5) Associate the video to be processed with the target data page, and write the playback source data of the video to be processed and the details on the target data page into the specified database.

[0078] Figure 5 A schematic diagram of the object storage device provided in an embodiment of the present invention is shown. Figure 5 As shown, the object storage device 500 includes: The first extraction module 501 is used to obtain data pages of multiple objects from the first site, each data page of an object includes detailed information of the object, perform feature extraction processing on each data page, and determine the first feature information corresponding to each data page; The second extraction module 502 is used to perform feature extraction processing on the playback source data of the object to be processed from at least one second site, and determine the second feature information of the object to be processed. The matching module 503 is used to determine the target data page that matches the object to be processed based on the first feature information and the second feature information; The database entry module 504 is used to associate the object to be processed with the target data page, and to write the playback source data of the object to be processed and the details of the target data page into a specified database.

[0079] The database entry module is used to associate the object to be processed with the target data page and write the playback source data of the object to be processed and the details of the target data page into the specified database.

[0080] Optionally, the inbound module is also used to bind the details information on the target data page as the descriptive text of the object to be processed.

[0081] Optionally, the device also includes a construction module for: building an index based on the first feature information corresponding to each data page; The matching module is also used to: search for target feature information that matches the second feature information in the first feature information based on the index, and use the data page corresponding to the target feature information as the target data page for matching the object to be processed.

[0082] Optionally, the first feature information includes multiple first terms, and the second feature information includes multiple second terms, wherein the multiple second terms include the name of the object; The matching module is also used to: determine the alias corresponding to the name of the object in the second feature information; search for candidate feature information that is the same as the name or alias of the object in the first feature information based on the index; calculate the first similarity between the first term and the corresponding second term for each first term in the candidate feature information, and take the sum of the first similarities between each first term and the corresponding second term in the candidate feature information as the second similarity between the candidate feature information and the second feature information; and take the candidate feature information with the second similarity greater than or equal to the first threshold as the target feature information similar to the second feature information.

[0083] Optionally, the device further includes a first filtering module for: performing a quality assessment on the data pages of multiple objects, determining a first score for each object's data page, and filtering the data pages of the multiple objects based on the first score.

[0084] Optionally, the profile page includes multiple detail fields and evaluation items. The multiple detail fields include at least one key field, and the evaluation items include one or more of the following: expected number of views, number of views already viewed, number of comments, and second rating. The number of detail fields is used to characterize the comprehensiveness of the profile page's content, and the number of key fields is used to characterize the usefulness of the profile page. The first filtering module is also used to: determine the first target interval to which the number of detail fields included in the data page belongs from multiple first intervals, determine the second target interval to which the number of key fields belongs from multiple second intervals, and determine the third target interval to which the evaluation items belong from multiple third intervals; calculate the weighted sum of the third score corresponding to the first target interval, the fourth score corresponding to the second target interval, and the fifth score corresponding to the third target interval, and use the weighted sum as the first score of the corresponding data page.

[0085] Optionally, the device further includes a second filtering module for: acquiring playback source data of multiple candidate objects from at least one second site; and filtering out the objects to be processed from the multiple candidate objects according to preset filtering rules.

[0086] Optionally, the filtering rules indicate one or more of the following: target type, target duration range, target publication time, target topic, and completeness threshold of playback source data; the second filtering module is further configured to: The step of selecting objects to be processed from the multiple candidate objects based on the playback source data of the multiple candidate objects and preset filtering rules includes: For each candidate object, based on the playback source data corresponding to the candidate object, determine one or more of the following: type, duration, release time, theme, and completeness of playback source data. From the plurality of candidate objects, select candidate objects that satisfy one or more of the following conditions, and use the candidate objects that satisfy one or more of the following conditions as objects to be processed: The type of the candidate object is consistent with the target type; The duration of the candidate is within the target duration range; The release time of the candidate object is consistent with the release time of the target object; The topic of the candidate object is consistent with the target topic; The completeness of the playback source data of the candidate object is greater than or equal to the completeness threshold.

[0087] The above-described apparatus can execute the method provided in the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the object storage method provided in the embodiments of the present invention.

[0088] Figure 6 A schematic diagram of the structure of an electronic device according to an embodiment of the present invention is shown. Figure 6 As shown, the electronic device includes: The system includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604. The processor 601, communication interface 602, and memory 603 communicate with each other via the communication bus 604. Memory 603 is used to store computer programs; When processor 601 executes a program stored in memory 603, it performs the following steps: Data pages of multiple objects are obtained from the first site. Each data page of an object includes detailed information about the object. Feature extraction processing is performed on each data page to determine the first feature information corresponding to each data page. Feature extraction processing is performed on the playback source data of the object to be processed from at least one second site to determine the second feature information of the object to be processed; Based on the first feature information and the second feature information, a target data page matching the object to be processed is determined; The object to be processed is associated with the target data page, and the playback source data of the object to be processed and the details on the target data page are written into the specified database.

[0089] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0090] The communication interface is used for communication between the aforementioned terminal and other devices.

[0091] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0092] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0093] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform any of the object loading methods described in the above embodiments.

[0094] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the object storage methods described in the above embodiments.

[0095] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).

[0096] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0097] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0098] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. An object warehousing method characterized by comprising: include: Retrieve information pages for multiple objects from the first site; each object's information page includes detailed information about the object. Perform feature extraction processing on each of the data pages to determine the first feature information corresponding to each of the data pages; Feature extraction processing is performed on the playback source data of the object to be processed from at least one second site to determine the second feature information of the object to be processed; Based on the first feature information and the second feature information, a target data page matching the object to be processed is determined; The object to be processed is associated with the target data page, and the playback source data of the object to be processed and the details on the target data page are written into the specified database.

2. The method according to claim 1, characterized in that, Associating the object to be processed with the target data page includes: Bind the detailed information on the target data page as the description text of the object to be processed.

3. The method according to claim 1 or 2, characterized in that, After determining the first feature information corresponding to each of the data pages, the method further includes: An index is created based on the first feature information corresponding to each of the aforementioned data pages; The step of determining the target data page matching the object to be processed based on the first feature information and the second feature information includes: Based on the index, the target feature information that matches the second feature information is found in the first feature information, and the data page corresponding to the target feature information is used as the target data page for matching the object to be processed.

4. The method of claim 3, wherein, The first feature information includes multiple first terms, and the second feature information includes multiple second terms, wherein the multiple second terms include the name of the object; The step of searching for target feature information that matches the second feature information in the first feature information based on the index includes: Determine the alias corresponding to the name of the object in the second feature information; Based on the index, search for candidate feature information that is the same as the name or alias of the object in the first feature information; For each first term in the candidate feature information, calculate the first similarity between the first term and the corresponding second term, and take the sum of the first similarities between each first term and the corresponding second term in the candidate feature information as the second similarity between the candidate feature information and the second feature information. Candidate feature information with a second similarity greater than or equal to the first threshold is taken as target feature information similar to the second feature information.

5. The method of claim 1, wherein, After retrieving data pages for multiple objects from the first site, the method further includes: The quality of the data pages of the multiple objects is evaluated to determine the first score of the data pages of each object. Based on the first score, the data pages of the multiple objects are filtered.

6. The method of claim 5, wherein, The data page includes multiple detail fields and evaluation items. The multiple detail fields include at least one key field, and the evaluation items include one or more of the following: expected number of views, number of views already viewed, number of comments, and second rating. The number of detail fields indicates the comprehensiveness of the information page's content, while the number of key fields indicates the usefulness of the information page. The quality assessment of the data pages of the multiple objects, and the determination of the first score for the data pages of each object, includes: For each object's data page, a first target interval is determined from multiple first intervals to which the number of detail fields included in the data page belongs, a second target interval is determined from multiple second intervals to which the number of key fields belongs, and a third target interval is determined from multiple third intervals to which the evaluation items belong; Calculate the weighted sum of the third score corresponding to the first target interval, the fourth score corresponding to the second target interval, and the fifth score corresponding to the third target interval, and use the weighted sum as the first score of the corresponding data page.

7. The method of claim 1, wherein, Before performing feature extraction processing on the playback source data of the object to be processed from at least one second site, the method further includes: Retrieve playback source data for multiple candidate objects from at least one secondary site; Based on the playback source data of the multiple candidate objects and the preset filtering rules, the objects to be processed are selected from the multiple candidate objects.

8. The method of claim 7, wherein, The filtering rules indicate one or more of the following: target type, target duration range, target release time, target theme, and completeness threshold of playback source data for the object to be processed. The step of selecting objects to be processed from the multiple candidate objects based on the playback source data of the multiple candidate objects and preset filtering rules includes: For each candidate object, based on the playback source data of the candidate object, determine one or more of the following: type, duration, release time, theme, and completeness of playback source data. From the plurality of candidate objects, select candidate objects that satisfy one or more of the following conditions, and use the candidate objects that satisfy one or more of the following conditions as objects to be processed: The type of the candidate object is consistent with the target type; The duration of the candidate is within the target duration range; The release time of the candidate object is consistent with the release time of the target object; The topic of the candidate object is consistent with the target topic; The completeness of the playback source data of the candidate object is greater than or equal to the completeness threshold.

9. An object warehousing device characterized by comprising: include: The first extraction module is used to obtain data pages of multiple objects from the first site. Each data page of an object includes detailed information of the object. The module performs feature extraction processing on each data page to determine the first feature information corresponding to each data page. The second extraction module is used to perform feature extraction processing on the playback source data of the object to be processed from at least one second site, and to determine the second feature information of the object to be processed. The matching module is used to determine the target data page that matches the object to be processed based on the first feature information and the second feature information; The database entry module is used to associate the object to be processed with the target data page, and to write the playback source data of the object to be processed and the details on the target data page into the specified database.

10. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method as described in any one of claims 1-8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-8.