Content error correction method, apparatus, device, and storage medium

By adding lifecycle markers to a pre-defined unified library and combining them with popularity rankings, the problem of low error correction accuracy in existing technologies has been solved, resulting in more accurate content error correction and a better user experience.

CN114239581BActive Publication Date: 2026-01-16HISENSE VISUAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111440044.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2026-01-16
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

Existing content correction methods cannot fully consider real-world application scenarios, resulting in low accuracy in error correction.

Method used

By adding lifecycle markers to a pre-defined unified database, and combining this with a popularity ranking list for clean word segmentation and matching, error correction methods are determined, including retaining or deleting unified content, and updating content based on factors such as lifecycle, popularity, and text similarity.

Benefits of technology

It improves the accuracy of content correction, ensures that search results are closer to user intent, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114239581B_ABST
    Figure CN114239581B_ABST
Patent Text Reader

Abstract

The application provides a content error correction method and device, equipment and a storage medium. The method obtains to-be-corrected uniform content in a preset uniform library, wherein the to-be-corrected uniform content includes error content and correct content corresponding to the error content. The life cycle of the to-be-corrected uniform content is determined according to a preset life cycle determination rule. The life cycle is compared with the current time. If the life cycle is less than the current time, the error correction mode of the to-be-corrected uniform content is determined according to the heat of the correct content. The preset uniform library is updated according to the error correction mode. The heat of the content is fully considered, so that the search result is closer to the user's intention, and the accuracy of content error correction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-computer interaction technology, and in particular to a content error correction method, apparatus, device and storage medium. Background Technology

[0002] With the continuous development of computer technology, users can access various online resources through search functions. For example, they can use the search function of video websites to find videos they want to watch, or use the search function of music software to find songs. During the process of numerous searches, inaccurate descriptions of the searched content may occur. For instance, if a TV series is titled "My Little Courtyard," and the user's search query is simply "My Little Courtyard," then the user's query needs to be corrected to display the search results for "My Little Courtyard" to the user.

[0003] Currently, for users' inaccurate statements, existing technologies typically combine dimensions such as pinyin or text similarity to match content in a pre-set database, thereby correcting errors.

[0004] However, existing error correction methods cannot fully consider real-world application scenarios, resulting in low accuracy in content error correction. Summary of the Invention

[0005] This application provides a content correction method, apparatus, device, and storage medium to solve the technical problem that existing error correction methods cannot fully consider practical application situations and have low accuracy in content correction.

[0006] Firstly, this application provides a content correction method, including:

[0007] Obtain the content to be corrected and normalized from the preset normalization library, wherein the content to be corrected and normalized includes the erroneous content and the correct content corresponding to the erroneous content;

[0008] The lifecycle of the content to be corrected and normalized is determined according to the preset lifecycle determination rules;

[0009] The lifecycle is compared with the current time. If the lifecycle is shorter than the current time, the correction method for the content to be corrected and normalized is determined based on the popularity of the correct content.

[0010] Update the preset normalization library according to the error correction method.

[0011] Herein, the embodiment of the present application provides a content correction method, aiming at the normalized content in a normalized library, the normalized content is marked by adding a life cycle, if the life cycle is less than the current time, it can be determined that the normalized content needs to be updated, in the correction process, the update of the content in the preset normalized library is combined with the heat, the heat of the content is fully considered, so that the search result is closer to the user's intention, and the accuracy of the content correction is improved.

[0012] The correction mode of the normalized content to be corrected is determined according to the heat of the correct content.

[0013] The correct content is cleaned, segmented and matched with entities in a preset heat list, and the correction mode is determined according to the processing result.

[0014] In the embodiment of the present application, the correct content in the normalized content and the content in the heat list can be cleaned, segmented and matched, so that the heat of the correct content is identified according to the content in the list, and the content with different heat can be processed.

[0015] In a possible design, the correction mode is determined according to the processing result.

[0016] If the first entity in the preset heat list matches the correct content, and the first entity has an associated entity, it is determined that the correction mode is to keep the normalized content to be corrected in the preset normalized library, and to add normalized content according to the associated entity in the preset normalized library, wherein the associated entity includes a pronunciation similar associated entity and a common name associated entity.

[0017] Herein, the embodiment of the present application matches the correct content of the normalized content with the heat list, and for the case that there is an entity matching the correct content in the heat list, and the entity has an associated entity, that is, the pronunciation is similar but at the same time has different heat content or the case that the series has an alias, the original normalized content can be saved in the preset normalized library, and the normalized content of the associated information is saved, when the user searches, the correct content and the associated information can be displayed at the same time, the search error caused by the alias and the pronunciation similar name when the user searches can be avoided, multiple information can be displayed to the user at the same time, the user can search for the content he wants according to the displayed different content (that is, the correct content and the associated information), and the user experience is further improved.

[0018] In a possible design, the correction mode is determined according to the processing result.

[0019] If the second entity matching the correct content exists in the preset heat list and the second entity has no associated entity, it is determined that the error correction mode is to keep the to-be-corrected normalized content in the preset normalization library.

[0020] Here, after the correct content of the normalized content is matched with the heat list, for the entity matching the correct content existing in the preset heat list and having no other term related to the correct content, it can be determined that the correct content of the normalized content is a hot broadcast or strong operation content, and the processing mode of the normalized content is to keep the normalized content, and the high heat term is kept in the preset normalization library. When the user searches, the high heat term can be displayed to the user according to the search content of the user, which further improves the accuracy of error correction and improves the user experience.

[0021] In a possible design, the determining the error correction mode according to the processing result comprises:

[0022] If the entity matching the correct content does not exist in the preset heat list, a product of a normalized heat of the correct content and a text similarity of the to-be-corrected normalized content is calculated, wherein the normalized heat is a ratio of a frequency of occurrence of the correct content in the preset normalization library to a total number of correct contents in the preset normalization library.

[0023] If the product is greater than a preset heat threshold, it is determined that the error correction mode is to keep the to-be-corrected normalized content in the preset normalization library.

[0024] If the product is not greater than the preset threshold, it is determined that the error correction mode is to delete the to-be-corrected normalized content in the preset normalization library.

[0025] Here, the embodiments of the present application update the normalized information in the preset normalization library in combination with the heat list, so that the preset normalization library can more accurately correct errors according to the search information of the user. For the correct content that cannot be matched with the entity in the heat list, the normalized heat of the correct content / text similarity (i.e., the text similarity of the normalized content) is first calculated, wherein the normalized heat of the correct content is the frequency of occurrence of the correct content / the total number of correct contents in the term library. If the formula result is too high, it indicates that the normalized term heat is high or the text similarity is high, and the normalized content is kept in the preset normalization library. If the formula result is low, it indicates that the normalized term heat is low or the text similarity is low, and the normalized content is deleted in the preset normalization library, so as to reduce the interference of the term without heat on the user search, so as to provide more accurate search results for the user according to the user search heat, further improve the accuracy of error correction, and further improve the user experience.

[0026] The calculating the product of the normalized heat of the correct content and the text similarity of the to-be-corrected normalized content comprises:

[0027] simulate, by a preset simulation interface, a request of deleting the to-be-corrected normalized content in the preset normalized library;

[0028] determine whether the correction result is accurate according to the simulation result;

[0029] if the correction result is not accurate, calculate a product of a normalized hotness of the correct content and a text similarity of the to-be-corrected normalized content.

[0030] Here, the embodiment of the present application first simulates the request of deleting the mapped content in the preset normalized library before calculating the product of the normalized hotness of the correct content and the text similarity of the to-be-corrected normalized content. If the correction result is accurate, the content does not need to be updated. If the result is not given, the content correction and update operation is continuously performed, thereby further saving system power consumption and improving correction efficiency.

[0031] In a possible design, the determining the correction mode for the to-be-corrected normalized content according to the hotness of the correct content includes:

[0032] obtaining a last update time of the to-be-corrected normalized content, and determining a time difference between a current time and the last update time;

[0033] if the time difference is greater than a preset update time threshold, determining the correction mode for the to-be-corrected normalized content according to the hotness of the correct content.

[0034] Here, the embodiment of the present application first obtains the time difference between the last update time and the current time of the normalized content before determining the correction mode for the normalized content to update the normalized content, thereby determining the update time length of the normalized content. If the time difference is greater than the preset update time threshold, it can be determined that the normalized content has not been updated for a period of time, and thus needs to be updated according to the hotness and other dimensional factors. Then, the correction mode for the to-be-corrected normalized content is determined according to the hotness of the correct content. Correspondingly, if the time difference is less than or equal to the preset update time threshold, it is determined that the normalized content has been updated for a period of time, and thus the normalized content does not need to be processed. This further reduces the content correction process, reduces system internal consumption, and also ensures the timeliness of the normalized content in the preset normalized library and increases the accuracy of the correction when the user searches.

[0035] In a possible design, the preset life cycle determination rule includes:

[0036] if the category of the to-be-corrected normalized content is a movie category, determining the life cycle of the to-be-corrected normalized content according to a movie release date and a first preset life cycle duration;

[0037] If the category of the to-be-error-corrected normalized content is a television series, a life cycle of the to-be-error-corrected normalized content is determined according to a television series broadcast end date and a second preset life duration.

[0038] If the category of the to-be-error-corrected normalized content is music, a life cycle of the to-be-error-corrected normalized content is determined according to a music ranking list ranking.

[0039] In the embodiments of the present application, different preset life cycle determination rules are set for different types of normalized content, so that an accurate life cycle reflecting the content heat of the normalized content can be determined according to different normalized content, so that the normalized content that has exceeded the life cycle, i.e., the content that has no heat at present, is deleted, and the normalized content that has not exceeded the life cycle, i.e., the normalized content that has heat at present, is screened and updated, the influence of the time factor on the heat of the normalized content is comprehensively considered, the content error correction is combined with the heat, the accuracy of the content error correction is further improved, and the user experience is further improved.

[0040] In a second aspect, the present application provides a content error correction device, comprising:

[0041] An acquisition module is configured to acquire to-be-error-corrected normalized content in a preset normalized library, wherein the to-be-error-corrected normalized content includes error content and correct content corresponding to the error content.

[0042] A first determination module is configured to determine a life cycle of the to-be-error-corrected normalized content according to a preset life cycle determination rule.

[0043] A second determination module is configured to compare the life cycle with a current time, and if the life cycle is less than the current time, determine an error correction mode for the to-be-error-corrected normalized content according to a heat of the correct content.

[0044] An update module is configured to update the preset normalized library according to the error correction mode.

[0045] In a possible design, the second determination module is specifically configured to:

[0046] The correct content is cleaned, segmented and matched with entities in a preset heat list, and an error correction mode is determined according to a processing result.

[0047] In a possible design, the second determination module is further specifically configured to:

[0048] If the first entity matching the correct content exists in the preset hot list, and the first entity has an associated entity, the error correction mode is determined as keeping the to-be-corrected normalized content in the preset normalized library, and adding normalized content according to the associated entity in the preset normalized library, wherein the associated entity includes a pronunciation similar associated entity and a common name associated entity.

[0049] In a possible design, the second determining module is specifically configured to:

[0050] If the second entity matching the correct content exists in the preset hot list, and the second entity has no associated entity, the error correction mode is determined as keeping the to-be-corrected normalized content in the preset normalized library.

[0051] In a possible design, the second determining module is specifically configured to:

[0052] If no entity matching the correct content exists in the preset hot list, a product of a normalized hot degree of the correct content and a text similarity of the to-be-corrected normalized content is calculated, wherein the normalized hot degree is a ratio of a frequency of occurrence of the correct content in the preset normalized library to a total number of correct contents in the preset normalized library.

[0053] If the product is greater than a preset hot threshold, the error correction mode is determined as keeping the to-be-corrected normalized content in the preset normalized library.

[0054] If the product is not greater than the preset threshold, the error correction mode is determined as deleting the to-be-corrected normalized content in the preset normalized library.

[0055] In a possible design, the second determining module is specifically configured to:

[0056] A request of deleting the to-be-corrected normalized content in the preset normalized library is simulated through a preset simulation interface.

[0057] According to the simulation result, it is determined whether the error correction result is accurate.

[0058] If the error correction result is not accurate, a product of a normalized hot degree of the correct content and a text similarity of the to-be-corrected normalized content is calculated.

[0059] In a possible design, the second determining module is specifically configured to:

[0060] A last update time of the to-be-corrected normalized content is acquired, and a time difference between a current time and the last update time is determined.

[0061] If the time difference is greater than the preset update time threshold, a correction mode for the to-be-corrected normalized content is determined according to the popularity of the correct content.

[0062] In a possible design, the first determining module is specifically configured to:

[0063] If the category of the to-be-corrected normalized content is a movie category, a life cycle of the to-be-corrected normalized content is determined according to a movie release date and a first preset life duration.

[0064] If the category of the to-be-corrected normalized content is a television series category, a life cycle of the to-be-corrected normalized content is determined according to a television series end date and a second preset life duration.

[0065] If the category of the to-be-corrected normalized content is a music category, a life cycle of the to-be-corrected normalized content is determined according to a music ranking.

[0066] In a third aspect, the present application provides a content correction device, comprising at least one processor and a memory.

[0067] The memory stores computer execution instructions.

[0068] The at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the content correction method as described in the first aspect and various possible designs of the first aspect.

[0069] In a fourth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores computer execution instructions, and when a processor executes the computer execution instructions, the content correction method as described in the first aspect and various possible designs of the first aspect is implemented.

[0070] In a fifth aspect, the present application provides a computer program product, comprising a computer program, and when a processor executes the computer program, the content correction method as described in the first aspect and various possible designs of the first aspect is implemented.

[0071] The content correction method, device, content correction device and storage medium provided by the present application, wherein the method is for the normalized content in the normalized library, the normalized content is marked by adding a life cycle, if the life cycle is less than the current time, it can be determined that the normalized content needs to be updated, in the correction process, the update of the content in the preset normalized library is combined with the popularity, the popularity of the content is fully considered, so that the search result is closer to the user's intention, and the accuracy of the content correction is improved. BRIEF DESCRIPTION OF DRAWINGS

[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the description of the embodiments or the prior art will be briefly introduced. Obviously, the accompanying drawings in the following description only represent some embodiments of the present application, and for those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0073] Figure 1 A schematic diagram of an application scenario of a content error correction method according to one or more embodiments of the present application;

[0074] Figure 2 A schematic diagram of a flow of a content error correction method provided by an embodiment of the present application;

[0075] Figure 3 A schematic diagram of a flow of another content error correction method provided by an embodiment of the present application;

[0076] Figure 4 A schematic diagram of a structure of a content error correction device provided by an embodiment of the present application;

[0077] Figure 5 A schematic diagram of a structure of a content error correction device provided by an embodiment of the present application.

[0078] The above-described drawings have shown the specific embodiments of the present disclosure, and the following detailed description will be given. These drawings and the following description are not intended to limit the scope of the present disclosure by any means, but to illustrate the concept of the present disclosure to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0079] The exemplary embodiments will be described in detail herein with reference to the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0080] The terms "first", "second", "third", "fourth" and the like in the description and in the claims of the present application, if any, are used for distinguishing between similar objects and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of these terms herein is merely used as a label to distinguish between the similar objects so that it implies no order or precedence. Also, the terms "comprises", "comprising", "includes", "including", "contains", "containing" and the like, if any, are intended to be inclusive and are used as the equivalent of "consisting essentially of". Accordingly, the use of these terms in the claims is not intended to be construed as a limitation on the scope of the application unless otherwise specifically stated.

[0081] Currently, users use voice search to search for popular TV shows, songs and other content more and more. As the names of videos, music and other content become more and more similar, after the user gives an inaccurate statement, the prior art usually combines the dimensions of pinyin or text similarity to match the content existing in the preset database, so as to realize error correction of the wrong content. Among them, the preset database (preset normalization library) includes normalized content, and the normalized content refers to directly mapping the wrong content to the correct content. For the wrong content given by the user, the correct content can be obtained by matching the wrong content in the preset normalization library and displayed to the user.

[0082] The error correction result given by the prior art is too random, which may not meet the user's demand at the first time, resulting in poor user experience. For example, for the popular TV series "My Small Courtyard", many users search for "My Small Courtyard". For the error correction library in the prior art, the search result of "Wu's Small Courtyard" (may give a movie many years ago) is given. However, considering the user context and the popularity of the movie, the user wants to search for the then popular "My Small Courtyard". In addition, there are some old content and new online movie content names that are very similar, which leads to the fact that if the user has a slight accent error, the given content is not what the user wants to search for. For example, during the broadcast of the popular TV series "Hello Jenny", the search volume of "Hello Jenny" (a movie released a few years ago) also soared. From the user context, it can be seen that the user wants to search for the TV series content. Or, the same movie content has different aliases, and the user searches for different names, but when searching for the old name, only some related news can be searched, and the probability of searching for the episode content is low.

[0083] To solve the above technical problems, the embodiment of the present application provides a content correction method and device, a content correction device and a storage medium. For the uniform content in the preset uniform library, the content name is marked by adding a life cycle, and through the cycle calculation of the life cycle, when the user gives a non-fully accurate statement, a search result closer to the user's idea is provided according to the multi-dimensional comprehensive calculation of the heat and the like.

[0084] Exemplarily, Figure 1 The application scenario of the content correction method according to one or more embodiments of the present application is shown in FIG. 1. Figure 1 As shown in the figure, the above architecture includes a user 100 and a content correction device 200.

[0085] The user 100 can interact with the content correction device 200 through the input device of the content correction device to send an intention to the content correction device 200. After receiving the intention, the content correction device 200 performs content correction according to the received intention through the internal processor, and can output the content correction result to the user through the output device to achieve the precise search purpose of the user.

[0086] In some embodiments, the content correction device 200 can be a display device such as a smart television, a personal computer, a smart refrigerator, or the like. Or a smart terminal such as a smart watch, a smart phone, or the like.

[0087] In some embodiments, the user 100 and the content correction device 200 can interact in the form of voice interaction, text interaction, or the like. For example, the content correction device 200 obtains the intention of the user through a microphone or a keyboard, a touch screen, or the like, and outputs the content correction result (search result) through a display screen or a loudspeaker, a horn, or the like.

[0088] It can be understood that the structure shown in the embodiment of the present application does not constitute a specific limitation on the architecture of the content correction system. In another possible implementation of the present application, the above architecture can include more or fewer components than the figure, or combine certain components, or split certain components, or different component arrangement, which can be determined according to the actual application scenario, and is not limited herein. Figure 1 The components shown in the figure can be realized in hardware, software, or a combination of software and hardware.

[0089] In addition, the network architecture and business scenario described in the embodiment of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. It can be known by those skilled in the art that, with the evolution of network architecture and the appearance of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0090] The technical solutions of the present application are described below with several embodiments as examples. For the same or similar concepts or processes, some embodiments may not be described again.

[0091] Figure 2 A flowchart of a content correction method provided by an embodiment of the present application is shown in the figure. The execution subject of the embodiment of the present application can be Figure 1 The content correction device 200 in the embodiment shown or the processor of the content correction device 200. The specific execution subject can be determined according to the actual application scenario. For example, Figure 2 The method includes the following steps:

[0092] S201: Obtain the to-be-corrected normalized content in a preset normalized library.

[0093] The to-be-corrected normalized content includes error content and correct content corresponding to the error content.

[0094] The normalized content included in the preset normalized library is determined by combining pinyin or text similarity in the prior art. The embodiment of the present application needs to correct the normalized content to update the preset normalized library, so that the normalized content in the preset normalized library is affected by the comprehensive heat and time factors, and accurate correction is realized for the error content given by the user.

[0095] S202: Determine the lifecycle of the to-be-corrected normalized content according to a preset lifecycle determination rule.

[0096] In one possible design, the preset lifecycle determination rule includes:

[0097] If the category of the to-be-corrected normalized content is movie, the lifecycle of the to-be-corrected normalized content is determined according to the movie release date and the first preset life duration (for example, lifecycle = movie release date + 90 days).

[0098] If the category of the to-be-corrected normalized content is television series, the lifecycle of the to-be-corrected normalized content is determined according to the end date of the television series and the second preset life duration (for example, lifecycle = end date of the television series + 90 days).

[0099] If the category of the to-be-corrected normalized content is music, the lifecycle of the to-be-corrected normalized content is determined according to the music chart ranking. Optionally, for music, threshold calculation is performed according to the chart ranking, and a threshold is assigned according to the hot song chart ranking.

[0100] It can be understood that the first preset life duration and the second preset life duration can be determined according to actual conditions, and the embodiment of the present application does not make specific limitations.

[0101] For other independent video content, event operation content, alias operation, etc., the lifecycle can be manually set by staff according to experience.

[0102] In a possible design, if the content type is not included in the above type range, the lifecycle is automatically empty.

[0103] Among them, the embodiment of the application sets different preset life cycle determination rules for different types of normalized content, so as to determine the life cycle that accurately reflects the content heat of different normalized content, so as to delete the normalized content that has exceeded the life cycle, that is, the content that has no heat at present, and screen and update the normalized content that has not exceeded the life cycle, that is, the normalized content that has heat at present, comprehensively consider the influence of time factor on the heat of normalized content, and combine the heat to correct the content, further improve the accuracy of content correction, and further improve the user experience.

[0104] In a possible design, after obtaining the life cycle of the normalized content, the following formula is calculated for the lifecycle with a value every day: current_time-lifecycle, if the calculation result is: current_time-lifecycle>update_time(set the time of lifecycle), then action is set to delete, and the normalized content is deleted in the preset normalized library; if current_time-lifecycle<update_time, the normalized content is retained, and the error correction mode is determined for the normalized content.

[0105] In the above manner, the effectiveness of the normalized content can be judged in advance, if the normalized content has exceeded the time limit (i.e. current_time-lifecycle>update_time), it is judged that the normalized content is invalid, if the normalized content has not exceeded the effective period (i.e. current_time-lifecycle<update_time), it is judged that the normalized content is valid, the normalized content is preliminarily screened according to the timeliness, and the interference of the word entry with long timeliness to the user search result is reduced.

[0106] S203: Compare the life cycle with the current time, if the life cycle is less than the current time, determine the error correction mode for the normalized content to be corrected according to the heat of the correct content.

[0107] In a possible design, for the life cycle of the normalized content, if greater than the current time (current_time), the life cycle of the normalized content is retained as lifecycle; if less than the current_time, the lifecycle is automatically empty.

[0108] According to the heat of the correct content, a correction mode for the normalized content to be corrected is determined, including: performing cleaning word segmentation and matching processing on the correct content and an entity in a preset heat list, and determining the correction mode according to a processing result.

[0109] In the embodiments of the present application, cleaning word segmentation and matching processing can be performed according to the correct content in the normalized content and the content in the heat list, so as to identify the heat of the correct content according to the content in the list, and facilitate processing of content with different heat.

[0110] In a possible design, according to the heat of the correct content, a correction mode for the normalized content to be corrected is determined, including: obtaining the last update time of the normalized content to be corrected, and determining the time difference between the current time and the last update time; if the time difference is greater than a preset update time threshold (for example, 365 days), the correction mode for the normalized content to be corrected is determined according to the heat of the correct content.

[0111] In the embodiments of the present application, the update time threshold can be determined according to actual conditions, and the present application does not make specific limitations.

[0112] In the embodiments of the present application, before determining the correction mode for the normalized content and updating the normalized content, the time difference between the last update time of the normalized content and the current time is first obtained, so as to determine the update time length of the normalized content. If the time difference is greater than the preset update time threshold, it can be determined that the normalized content has not been updated for a period of time, and therefore needs to be updated according to the heat and other dimensional factors. Then, the correction mode for the normalized content to be corrected is determined according to the heat of the correct content. Correspondingly, if the time difference is less than or equal to the preset update time threshold, it is determined that the normalized content has been updated for a period of time, and the normalized content does not need to be processed, further reducing the content correction process and reducing the system consumption, and ensuring the timeliness of the normalized content in the preset normalized library and increasing the accuracy of the correction when the user searches.

[0113] S204: updating the preset normalized library according to the correction mode.

[0114] Herein, the embodiment of the present application provides a content error correction method, for the normalized content in a normalized library, the normalized content is marked by adding a life cycle, if the life cycle is less than the current time, it can be determined that the normalized content needs to be updated, in the error correction process, the update of the content in the preset normalized library is combined with the heat, the heat of the content is fully considered, so that the search result is closer to the user's intention, and the accuracy of the content error correction is improved.

[0115] In some embodiments, for the normalized content of different heat types, the embodiment of the present application determines different error correction modes, and the specific mode is as follows:

[0116] In a possible design, the error correction mode is determined according to the processing result, including: if there is a first entity matching the correct content in the preset heat list, and the first entity has an associated entity, it is determined that the error correction mode is to keep the normalized content to be corrected in the preset normalized library, and to add the normalized content according to the associated entity in the preset normalized library, wherein the associated entity includes a pronunciation similar associated entity and a common name associated entity.

[0117] Herein, after the correct content of the normalized content is matched with the heat list, for the case that there is an entity matching the correct content in the heat list, and the entity has an associated entity, that is, the pronunciation is similar but at the same time has different heat content or the case that the episode has an alias, the original normalized content can be saved in the preset normalized library, and the normalized content of the associated information is saved, when the user searches, the correct content and the associated information can be displayed at the same time, the search error situation caused by the alias and the pronunciation similar name in the user search is avoided, multiple information can be displayed to the user at the same time, the user can search for the content that the user wants according to the displayed different content (that is, the correct content and the associated information), and the user experience is further improved.

[0118] In a possible design, the error correction mode is determined according to the processing result, including: if there is a second entity matching the correct content in the preset heat list, and the second entity has no associated entity, it is determined that the error correction mode is to keep the normalized content to be corrected in the preset normalized library.

[0119] Herein, after the correct content of the normalized content is matched with the heat list, for the case that there is an entity matching the correct content in the preset heat list, and there is no other term related to the correct content, it can be determined that the correct content of the normalized content is a hot broadcast or a strong operation content, and then the processing mode of the normalized content is to keep the normalized content, and the high-heat term is kept in the preset normalized library, when the user searches, the high-heat term can be displayed to the user according to the search content of the user, the accuracy of the error correction is further improved, and the user experience is improved.

[0120] In a possible design, the error correction manner is determined according to the processing result, including: if there is no entity in the preset hot list that matches the correct content, calculating a product of a normalized hotness of the correct content and a text similarity between the normalized content to be corrected and the correct content, where the normalized hotness is a ratio of a frequency of occurrence of the correct content in a preset normalized library to a total number of correct contents in the preset normalized library.

[0121] If the product is greater than a preset hotness threshold, it is determined that the error correction manner is to keep the normalized content to be corrected in the preset normalized library; if the product is not greater than the preset threshold, it is determined that the error correction manner is to delete the normalized content to be corrected in the preset normalized library.

[0122] It can be understood that the preset threshold can be determined according to actual conditions, and the embodiments of the present application do not make specific limitations.

[0123] Optionally, the text similarity is obtained by using a Smooth Inverse Frequency algorithm.

[0124] The product of the normalized hotness of the correct content and the text similarity of the normalized content to be corrected includes: simulating, through a preset simulation interface, a request of deleting the normalized content to be corrected in the preset normalized library; determining, according to a simulation result, whether the error correction result is accurate; if the error correction result is not accurate, calculating the product of the normalized hotness of the correct content and the text similarity of the normalized content to be corrected.

[0125] Here, the embodiments of the present application first simulate the deletion of the mapped request in the pre-normalized library before calculating the product of the normalized hotness of the correct content and the text similarity of the normalized content to be corrected, if the error correction result is accurate, there is no need to update the content, if the result is not given, the content error correction update operation is continuously performed, further saving the system power consumption and improving the error correction efficiency.

[0126] Here, the embodiments of the present application update the normalized information in the preset normalized library in combination with the hot list, so that the preset normalized library can more accurately correct errors according to the search information of the user, wherein for the correct content that cannot be matched with the entity in the hot list, first calculate the normalized hotness of the correct content / text similarity (i.e. the text similarity of the normalized content), where the normalized hotness of the correct content is the frequency of occurrence of the correct content / the total number of correct contents in the word library, if the formula result is too high, it means that the normalized word hotness is high or the text similarity is high, then keep the normalized content in the preset normalized library, if the formula result is too low, it means that the normalized word hotness is weak or the text similarity is low, then delete the normalized content in the preset normalized library, so as to reduce the interference of the word without hotness on the user search, so as to provide more accurate search results for the user according to the user search hotness, further improving the accuracy of error correction and further improving the user experience.

[0127] In some embodiments, the embodiments of the present application realize the differentiation of content popularity by setting a flag value for the normalized content, and the flag value is 0 by default. Figure 3 The flowchart of another content correction method provided by the embodiments of the present application is shown in FIG. 3, and the method includes the following steps. Figure 3

[0128] S31: Obtain the last update time of the normalized content to be corrected, and determine the time difference between the current time and the last update time.

[0129] S32: Is the time difference greater than the preset update time threshold?

[0130] No, perform S321: do not update the preset normalized library, and end.

[0131] Yes, perform S322: Is there an entity matching the correct content in the preset popularity list?

[0132] If the result of S322 is no, perform S3221: flag=0; and then perform S3222: Is the product of the normalized popularity of the correct content and the text similarity of the normalized content to be corrected greater than the preset popularity threshold? If the result of S3222 is yes, perform S32221: flag=2, and delete the normalized content in the preset normalized library; if the result of S3222 is no, perform S32222: flag=0, and keep the normalized content to be corrected in the preset normalized library.

[0133] If the result of S322 is yes, perform S3220: Is there an associated entity matching the normalized content? If the result of S3220 is no, perform S32201: flag=1, and determine the correction mode as keeping the normalized content to be corrected in the preset normalized library. If the result of S3220 is yes, perform S32202: if the associated entity is a pronunciation similar associated entity, flag=3; if the associated entity is a common name associated entity, flag=4. Keep the normalized content to be corrected in the preset normalized library, and add the normalized content according to the associated entity in the preset normalized library.

[0134] From the above steps, it can be known that flag=0 is the content to be processed; flag=1 is the hot broadcast or strong operation content, which is not processed; flag=2, perform action=delete operation; flag=3, for the case of pronunciation similar but different popularity content, search the search content of different popularity words at the same time; flag=4, for the case of episodes and other alias, search the search content of multiple names at the same time, and do not participate in the following loop (the flag value is not set, which can be defaulted to 0).

[0135] Figure 4 ​A structural schematic diagram of a content correction device provided by an embodiment of the present application is shown in Figure 4 The device of the embodiment of the present application includes an acquisition module 401, a first determination module 402, a second determination module 403 and an update module 404. The content correction device herein can be the content correction device 200 itself, or a chip or integrated circuit implementing the functions of the content correction device 200. It should be noted that the division of the acquisition module 401, the first determination module 402, the second determination module 403 and the update module 404 is only a logical division, and physically, they can be integrated or independent.

[0136] The acquisition module is configured to acquire, in a preset normalization library, to-be-corrected normalized content, wherein the to-be-corrected normalized content includes error content and correct content corresponding to the error content.

[0137] The first determination module is configured to determine, according to a preset life cycle determination rule, a life cycle of the to-be-corrected normalized content.

[0138] The second determination module is configured to compare the life cycle with a current time, and if the life cycle is less than the current time, determine, according to a heat of the correct content, a correction mode for the to-be-corrected normalized content.

[0139] The update module is configured to update the preset normalization library according to the correction mode.

[0140] In a possible design, the second determination module is specifically configured to:

[0141] perform cleaning, word segmentation and matching processing on the correct content and entities in a preset heat list, and determine the correction mode according to a processing result.

[0142] In a possible design, the second determination module is specifically configured to:

[0143] If there is a first entity in the preset heat list that matches the correct content, and the first entity has an associated entity, it is determined that the correction mode is to keep the to-be-corrected normalized content in the preset normalization library, and to add normalized content according to the associated entity in the preset normalization library, wherein the associated entity includes a pronunciation similar associated entity and a common name associated entity.

[0144] In a possible design, the second determination module is specifically configured to:

[0145] If there is a second entity in the preset heat list that matches the correct content, and the second entity has no associated entity, it is determined that the correction mode is to keep the to-be-corrected normalized content in the preset normalization library.

[0146] In a possible design, the second determination module is specifically configured to:

[0147] if the entity matching the correct content does not exist in the preset hot list, calculating a product of a normalized hotness of the correct content and a text similarity between the correct content and the normalized content to be corrected, wherein the normalized hotness is a ratio of a frequency of occurrence of the correct content in a preset normalized library to a total number of correct contents in the preset normalized library;

[0148] if the product is greater than a preset hotness threshold, determining that the correction mode is to keep the normalized content to be corrected in the preset normalized library;

[0149] if the product is not greater than the preset threshold, determining that the correction mode is to delete the normalized content to be corrected in the preset normalized library.

[0150] In a possible design, the second determining module is specifically configured to:

[0151] simulate, through a preset simulation interface, a request of deleting the normalized content to be corrected in the preset normalized library;

[0152] determine, according to a simulation result, whether the correction result is accurate;

[0153] if the correction result is not accurate, calculating a product of a normalized hotness of the correct content and a text similarity between the correct content and the normalized content to be corrected.

[0154] In a possible design, the second determining module is specifically configured to:

[0155] obtain a last update time of the normalized content to be corrected, and determine a time difference between a current time and the last update time;

[0156] if the time difference is greater than a preset update time threshold, determining, according to a hotness of the correct content, a correction mode for the normalized content to be corrected.

[0157] In a possible design, the first determining module is specifically configured to:

[0158] if the category of the normalized content to be corrected is a movie category, determining a life cycle of the normalized content to be corrected according to a movie release date and a first preset life duration;

[0159] if the category of the normalized content to be corrected is a television series category, determining a life cycle of the normalized content to be corrected according to a television series end date and a second preset life duration;

[0160] if the category of the normalized content to be corrected is a music category, determining a life cycle of the normalized content to be corrected according to a music ranking list ranking.

[0161] Figure 5 A content correction device (which can be a server or a terminal) provided by the embodiments of the present application Figure 1FIG. 1 is a structural schematic diagram of a content correction device 200 according to an embodiment of the present application. The components shown herein, their connections and relationships, and their functions are merely examples and do not limit the implementation of the present application described and / or claimed herein.

[0162] As shown in FIG. 1, the content correction device includes a processor 501 and a memory 502, each of which is connected to each other by different buses and can be installed on a common mainboard or otherwise as needed. The processor 501 can process instructions executed within the terminal, including instructions stored in the memory or graphical information stored on the memory to be displayed on external input / output devices such as display devices coupled to the interface. In other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memories as needed. Figure 5 Figure 5 In FIG. 1, the processor 501 is taken as an example.

[0163] The memory 502, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs and modules, such as program instructions / modules corresponding to the method of the content correction device in the embodiments of the present application (for example, the acquisition module 401, the first determination module 402, the second determination module 403 and the update module 404 shown in FIG. 1). Figure 4 The processor 501 performs various functional applications and data processing of the content correction device by running the non-transitory software programs, instructions and modules stored in the memory 502, that is, implements the method of the content correction device in the method embodiments described above.

[0164] The content correction device can also include an input device 503 and an output device 504. The processor 501, the memory 502, the input device 503 and the output device 504 can be connected by a bus or other means, Figure 5 In FIG. 1, the connection by the bus is taken as an example.

[0165] The input device 503 can receive input digital or character information and generate key signal inputs related to user settings and function controls of the content correction device, such as touch screens, keypads, mice, or multiple mouse buttons, trackballs, joysticks, and the like. The output device 504 can be a display device of the content correction device and the like. The display device can include but is not limited to liquid crystal displays (LCDs), light-emitting diode (LED) displays, and plasma displays. In some embodiments, the display device can be a touch screen.

[0166] The content correction device of the embodiments of the present application can be used to execute the technical solutions in the method embodiments described above, which have similar implementation principles and technical effects, and will not be described here.​

[0167] The embodiment of the present application further provides a computer readable storage medium, which stores computer execution instructions, and the computer execution instructions are used for realizing the content correction method when executed by a processor.

[0168] The embodiment of the present application further provides a computer program product, which comprises a computer program, and the computer program is used for realizing the content correction method when executed by a processor.

[0169] In several embodiments provided by the present application, it should be understood that the disclosed system, device and method can be implemented by other ways. For example, the device embodiments described above are only schematic, and for example, the division of the units is only a logical function division, and there can be another division way in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0170] In addition, each function unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist alone physically, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware, or in the form of a software function unit.

[0171] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. The application is intended to cover any variations, uses or adaptations of the application following, in general, the principles of the application and including such departures from the present disclosure that come within known

[0172] It should be understood that the present disclosure is not limited to the precise structures as has been described and illustrated above, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A content error correction method, characterized by, The method comprises the following steps: obtaining to-be-corrected content in a preset normalization library, wherein the to-be-corrected content comprises error content and correct content corresponding to the error content; determining a life cycle of the to-be-corrected content according to a preset life cycle determination rule; comparing the life cycle with a current time, and if the life cycle is less than the current time, determining a correction mode for the to-be-corrected content according to a heat of the correct content; updating the preset normalization library according to the correction mode; the preset life cycle determination rule comprises: if the category of the to-be-corrected content is a movie, determining the life cycle of the to-be-corrected content according to a movie release date and a first preset life duration; if the category of the to-be-corrected content is a TV series, determining the life cycle of the to-be-corrected content according to a TV series end date and a second preset life duration; if the category of the to-be-corrected content is music, determining the life cycle of the to-be-corrected content according to a music ranking; the determining of the correction mode for the to-be-corrected content according to the heat of the correct content comprises: obtaining a last update time of the to-be-corrected content, and determining a time difference between the current time and the last update time; if the time difference is greater than a preset update time threshold, performing cleaning word segmentation and matching processing on the correct content and an entity in a preset heat list, and determining a correction mode according to a processing result.

2. The method of claim 1, wherein, the determining of the correction mode according to the processing result comprises: if there is a first entity matching the correct content in the preset heat list, and the first entity has an associated entity, determining that the correction mode is to keep the to-be-corrected content in the preset normalization library, and adding a normalization content according to the associated entity in the preset normalization library, wherein the associated entity comprises a pronunciation similar associated entity and a common name associated entity.

3. The method of claim 1, wherein, the determining of the correction mode according to the processing result comprises: if there is a second entity matching the correct content in the preset heat list, and the second entity has no associated entity, determining that the correction mode is to keep the to-be-corrected content in the preset normalization library.

4. The method of claim 1, wherein, the determining of the correction mode according to the processing result comprises: if there is no entity matching the correct content in the preset heat list, calculating a product of a normalization heat of the correct content and a text similarity of the to-be-corrected content, wherein the normalization heat is a ratio of a frequency of the correct content in the preset normalization library to a total number of correct contents in the preset normalization library; if the product is greater than a preset heat threshold, determining that the correction mode is to keep the to-be-corrected content in the preset normalization library; if the product is not greater than the preset threshold, determining that the correction mode is to delete the to-be-corrected content in the preset normalization library.

5. The method of claim 4, wherein, the calculating of the product of the normalization heat of the correct content and the text similarity of the to-be-corrected content comprises: simulating a request for deleting the to-be-corrected content in the preset normalization library through a preset simulation interface. According to the simulation result, it is determined whether the error correction result is accurate or not; If the error correction result is not accurate, the product of the normalized heat degree of the correct content and the text similarity of the content to be error corrected is calculated.

6. A content error correction apparatus characterized by comprising: Comprise: An acquisition module is configured to acquire content to be error corrected in a preset normalization library, wherein the content to be error corrected comprises error content and correct content corresponding to the error content; A first determination module is configured to determine a life cycle of the content to be error corrected according to a preset life cycle determination rule; A second determination module is configured to compare the life cycle with a current time, and if the life cycle is less than the current time, determine an error correction mode for the content to be error corrected according to a heat degree of the correct content; An update module is configured to update the preset normalization library according to the error correction mode; The preset life cycle determination rule comprises: if a category of the content to be error corrected is a movie category, determining the life cycle of the content to be error corrected according to a movie release date and a first preset life duration; If the category of the content to be error corrected is a TV series category, determining the life cycle of the content to be error corrected according to a TV series end date and a second preset life duration; If the category of the content to be error corrected is a music category, determining the life cycle of the content to be error corrected according to a music ranking. The second determination module is specifically configured to acquire a last update time of the content to be error corrected, and determine a time difference between the current time and the last update time; If the time difference is greater than a preset update time threshold, the correct content is cleaned, segmented and matched with entities in a preset heat degree list, and the error correction mode is determined according to the processing result.

7. A content error correction apparatus characterized by comprising: Comprise: At least one processor; And The memory is in communication connection with the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by the processor to implement the content error correction method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Information processing method and device

    CN104216995A

  • Query processing method and device and computer readable storage medium

    CN111061750A

  • Pre-push content management method and device and computer equipment

    CN112241413A