Document validity determination method and device and computer equipment
By combining document content, timestamps, and user data to evaluate document validity, the problem of low accuracy in existing technologies is solved, enabling more accurate document management and knowledge base updates.
Patent Information
- Application Number
- CN202511094711.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-21
AI Technical Summary
In existing technologies, determining document validity by manually annotating document version numbers and timestamps has low accuracy and fails to consider the actual usability of the document content.
By combining multimodal data such as document content, timestamps, user access data, and comment data, the effectiveness of documents is evaluated through a pre-set model, including the calculation of sentiment score, time decay, and access popularity, and a knowledge graph is generated to manage documents.
It improves the accuracy of document validity determination, enabling more accurate identification of document practical value and timeliness, and enhances the efficiency of knowledge base management.
Smart Images

Figure CN120996028A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to a method and device for determining document validity and a computer device. BACKGROUND
[0002] A platform knowledge base usually stores a large number of documents, for example, technical documents, fault solution documents, etc., which can be consulted and learned by staff.
[0003] The documents in the knowledge base are updated according to the iteration of technology or product, etc. The document version before updating, the document version after updating, etc. are stored in the knowledge base. In order to identify the validity of the document, the current scheme is to distinguish the validity of different documents by manually labeling tags (such as version number, timestamp, etc.). The accuracy of the document validity determined by this method is low. SUMMARY
[0004] In order to solve the above technical problems, the present disclosure provides a method and device for determining document validity and a computer device.
[0005] In a first aspect, the present disclosure provides a method for determining document validity, which comprises:
[0006] Obtaining target information, the target information comprising at least two of the following: part or all of the content included in the first document, the timestamp of the first document, the access data of the user to the first document, and the comment data of the user to the first document; determining the validity of the first document based on the target information.
[0007] In some optional embodiments, determining the validity of the first document based on the target information comprises:
[0008] Inputting part or all of the content included in the first document, the access data of the user to the first document, and the comment data of the user to the first document into a preset model, outputting a sentiment score of the first document through the preset model; determining the validity of the first document based on the sentiment score, the timestamp of the first document, and the access data of the user to the first document.
[0009] In some optional embodiments, the timestamp of the first document comprises a publishing time;
[0010] Determining the validity of the first document based on the sentiment score, the timestamp of the first document, and the access data of the user to the first document comprises:
[0011] determine an age of the first document as a difference between the current time and the publishing time; determine a time decay amount based on the age and an access heat of the first document, the access heat being determined based on access data of the first document and access data of a second document, the second document being a document stored in the storage space and different from the first document; determine the validity of the first document based on the sentiment score and the time decay amount.
[0012] In some optional embodiments, determining the validity of the first document based on the sentiment score and the time decay amount comprises:
[0013] determining the validity of the first document based on the sentiment score, the time decay amount and the age of the first document.
[0014] In some optional embodiments, determining the validity of the first document based on the sentiment score, the time decay amount and the age of the first document comprises:
[0015] determining the validity of the first document based on the sentiment score, a weight corresponding to the sentiment score, the time decay amount, a weight corresponding to the time decay amount and the age of the first document.
[0016] In some optional embodiments, the weight corresponding to the sentiment score and the weight corresponding to the time decay amount are determined based on a type of the first document.
[0017] In some optional embodiments, the method further comprises:
[0018] generating a knowledge graph, wherein the knowledge graph comprises an association relationship between documents and attribute information of each document, the attribute information of each document comprising at least one of a timeliness index, a decay factor and a document type, wherein the timeliness index is determined based on a sentiment score of the document and a time decay amount of the document, and the decay factor is determined based on an age of the document.
[0019] In some optional embodiments, the condition for generating the knowledge graph comprises at least one of:
[0020] detecting a new document published by a user; detecting that at least one of the timeliness index and the decay factor of the document does not meet a preset condition; meeting a preset period for updating the knowledge graph.
[0021] In a second aspect, the present disclosure provides a device for determining validity of a document, the device comprising:
[0022] an obtaining module configured to obtain target information, the target information comprising at least two of: part or all of content included in a first document, a timestamp of the first document, access data of the first document by a user, and comment data of the first document by the user; and a determining module configured to determine validity of the first document based on the target information.
[0023] In a third aspect, the present disclosure provides a computer device, comprising:
[0024] The memory and the processor are connected with each other in communication, the memory stores computer instructions, and the processor executes the computer instructions to perform the method of the first aspect and any of the corresponding embodiments.
[0025] Compared with the prior art, the technical scheme provided by the embodiments of the present disclosure has the following advantages:
[0026] The document validity determination method provided by the present embodiment can improve the accuracy of determining the validity of the document by combining the content, timestamp, user access data, comment data and other multi-modal data of the document to evaluate the validity of the document. BRIEF DESCRIPTION OF DRAWINGS
[0027] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the drawings required to be used in the embodiments or the prior art description will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0029] Figure 1 A flowchart of a document validity determination method provided by the embodiments of the present disclosure is shown in the figure.
[0030] Figure 2 A schematic diagram of a knowledge graph corresponding to a graph data structure provided by the embodiments of the present disclosure is shown in the figure.
[0031] Figure 3 A schematic diagram of a knowledge graph corresponding to a tree data structure provided by the embodiments of the present disclosure is shown in the figure.
[0032] Figure 4 A structure connection diagram of a document validity determination device provided by the embodiments of the present disclosure is shown in the figure.
[0033] Figure 5 A structure connection diagram of a computer device provided by the embodiments of the present disclosure is shown in the figure. DETAILED DESCRIPTION
[0034] In order to more clearly illustrate the above-mentioned purposes, features and advantages of the present disclosure, the schemes of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0035] Many particular details are set forth in the following description in order to provide a thorough understanding of the present disclosure. However, the present disclosure can be practiced according to the claims without some or all of these details. Indeed, some embodiments of the present disclosure can be specifically adapted to the needs and circumstances of a specific embodiment or implementation.
[0036] Before the embodiments of the present disclosure are described in detail, the technical background related to the present disclosure is described here, so that those skilled in the art can have a clearer understanding of the embodiments of the present disclosure.
[0037] A large number of documents are usually stored in the knowledge base of the enterprise digital management platform, and these documents are continuously updated according to the iteration of enterprises, technologies or products, so that the staff can check and learn when needed. For example, in the knowledge base of a cloud vendor, the stored documents include, but are not limited to, application program interface documents, fault solution documents, etc. In order to identify the effectiveness of the documents, in one possible solution, the effectiveness of the documents can be manually annotated based on the version number, timestamp, etc. of the documents. When the staff checks the documents, the effectiveness of the documents is judged according to the foregoing annotation, and the documents with higher effectiveness are preferentially selected for access. When the management personnel clean up the documents in the knowledge base, the effectiveness of the documents is also judged according to the manual annotation label, and the documents with lower effectiveness are preferentially selected for cleaning.
[0038] The above-mentioned method has strong subjectivity, and only distinguishes the effectiveness of the documents based on a single dimension such as the version number, timestamp, etc. of the documents, without considering the actual use value of the document content, so the accuracy of the determined effectiveness of the documents is low and the user is easily misled.
[0039] Based on this, the embodiments of the present disclosure provide a method for determining the effectiveness of a document, which can combine the content, timestamp, access data of the document by the user, comment data and other multi-modal data to evaluate the effectiveness of the document, thereby improving the accuracy of the determined effectiveness of the document.
[0040] According to the embodiments of the present disclosure, a method for determining the effectiveness of a document is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described here can be executed in an order different from that shown here.
[0041] Reference Figure 1FIG. 1 is a flowchart of a method for determining the validity of a document according to an embodiment of the present disclosure. The method can be applied to a computer device, or to a module or chip in the computer device. The computer device can be a single computer or an integration of multiple computers. As shown in FIG. 1, the method includes the following steps: Figure 1
[0042] S101, obtaining target information.
[0043] The target information includes at least two of the following: part or all of the content included in the first document, a timestamp of the first document, access data of the first document by a user, and comment data of the first document by the user.
[0044] For example, the timestamp of the first document can include, but is not limited to, at least one of the following: a publication time of the first document, and an update time of the first document.
[0045] For example, the access data of the first document by the user can include, but is not limited to, at least one of the following: access records of the first document by the user, and operation records of the first document by the user. For example, the access records of the first document by the user can include, but are not limited to, at least one of the following: a user identifier (such as a user ID), an access time, and an access duration. The operation records of the first document by the user can include, but are not limited to, records of various operations performed by the user on the first document, such as collection, download, and sharing.
[0046] Optionally, the access records of the first document by the user can be obtained from a document access log.
[0047] For example, the comment data of the first document by the user can include, but is not limited to, at least one of the following: comment content, a score, and a timestamp of publishing the comment.
[0048] Optionally, the target information can further include a file identifier (such as a file ID) and a label of the file. For example, the label of the file can include, but is not limited to, at least one of the following: a document type (such as an application program interface document and an operation and maintenance scheme document), a document title, and a document version number.
[0049] Optionally, the first document can be any document stored in a knowledge base.
[0050] Optionally, the part or all of the content included in the first document, the timestamp of the first document, the access data of the first document by the user, and the comment data of the first document by the user can be obtained from one or more storage locations, which are not limited in the embodiments of the present disclosure.
[0051] Optionally, the target information can be original information obtained, or information obtained after pre-processing of the original information. Exemplarily, the pre-processing includes, but is not limited to, converting a data format, performing filtering of valid data (or described as removing invalid data), and the like. Exemplarily, the converting of the data format can be, for example, converting a timestamp into a unified standard format (such as a Unix format or other formats, etc.). The filtering of the valid data can be, for example, deleting blank comments published by a user, deleting abnormal access records of a user, and the like.
[0052] S102, determining validity of the first document based on the target information.
[0053] based on Figure 1 According to the technical solutions shown, the validity of the document can be evaluated by combining the content of the document, the timestamp, the access data of the user to the document, the comment data and the like, so that the accuracy of the determined validity of the document can be improved.
[0054] The specific implementation of determining the validity of the first document based on the target information will be introduced below.
[0055] In some embodiments, the validity of the first document can be determined in combination with a sentiment score of the first document. The sentiment score is used to represent the practical value of the document, and can also be understood as a score of the content of the document in the dimension of practicality.
[0056] Specifically, part or all of the content included in the first document, the access data of the user to the first document and the comment data of the user to the first document can be input into a preset model, and the sentiment score of the first document can be output by the preset model.
[0057] Exemplarily, the preset model can be a natural language processing (NLP) model. Of course, the preset model can also be other machine learning models or deep learning models.
[0058] Optionally, in the embodiments of the present application, the value range of the sentiment score of the document can be greater than or equal to -1 and less than or equal to 1.
[0059] Optionally, to further improve the accuracy of the determination result of the validity of the document, the sentiment score of the document can also be normalized. The value range of the normalized sentiment score is greater than or equal to 0 and less than or equal to 1.
[0060] Exemplarily, the sentiment score can be normalized by using the following formula 1.
[0061]
[0062] In formula 1, sentiment score is the sentiment score of the first document, and max sentiment score is the maximum value of the sentiment score of the first document.index S is the normalized sentiment score, and S is the sentiment score before normalization.
[0063] Further, the validity of the first document can be determined based on the sentiment score, the timestamp of the first document, and the access data of the first document by the user.
[0064] Optionally, in determining the validity of the first document based on the sentiment score, the timestamp of the first document, and the access data of the first document by the user, a time decay amount can be determined based on the timestamp of the first document and the access data of the first document by the user first, and then the validity of the first document can be determined based on the time decay amount and the sentiment score.
[0065] As an implementation of determining the time decay amount, the timestamp of the first document can include a publication time, and a difference between the current time and the publication time can be determined as the age of the first document; the time decay amount can be determined based on the age of the first document and the access popularity of the first document.
[0066] The access popularity of the first document will be introduced first.
[0067] The access popularity of the first document can be determined based on the access data of the first document and the access data of the second document, and the second document is a document stored in the storage space other than the first document. Optionally, the second document can include one or more documents. For example, the storage space can be a knowledge base of each platform, and the first document and the second document can constitute all documents stored in the knowledge base.
[0068] In some implementations, the access popularity of a document (such as the first document, the second document, etc.) can be determined based on the access amount of the document. The access amount of the document can be determined based on the access data of the document by the user. Optionally, the access amount of the document can be the access amount within a preset time period, such as the last 30 days, the last week, the last 3 months, etc.
[0069] Optionally, the access amount of the first document can be determined based on the access data of the first document. The maximum access amount and the minimum access amount in the access amount of all documents can be determined based on the access data of the first document and the access data of the second document. The access popularity of the first document can be determined based on the access amount of the first document, the maximum access amount, and the minimum access amount.
[0070] In some implementations, the access data of the document includes the access times of the document, and the access amount of the document can be determined based on the access times of the document.
[0071] For example, the access times can include, but are not limited to, the number of one or more operations such as collection, download, sharing, and access.
[0072] As a concrete example, the access popularity of the first document can be determined based on formulas 2, 3, and 4.
[0073] The number of visits to the first document can be determined based on Formula 2.
[0074] current activity =x·visits+y·favorites+z·shares Formula 2
[0075] Where, current activity Here, x represents the number of visits to the first document, y represents the number of times the first document has been viewed, y represents the number of times the first document has been favorited and / or downloaded, and z represents the number of times the first document has been shared. x is the weight of the number of visits, y is the weight of the number of favorites and / or downloads, and z is the weight of the number of shares. For example, x, y, and z can be set to 1.0, 1.5, and 2, respectively. The values of x, y, and z can be determined based on the actual business scenario.
[0076] Similarly, after obtaining the number of visits to the first document, it can be normalized using Formula 3, mapping the number of visits to the [0,1] interval. The normalized number of visits to the first document is then determined as its access popularity. The access popularity of the first document can be determined based on Formula 3.
[0077]
[0078] Among them, normalized behavior The normalized number of visits to the first document, i.e., the access popularity of the first document, has a value range of [0,1]. activity max represents the number of views for the first document. activity For the maximum number of visits, min activity Minimum number of visits.
[0079] Based on Formula 3 above, it can be seen that when the number of visits to the first document approaches the minimum number of visits, the access popularity of the first document approaches 0. When the number of visits to the first document approaches the maximum number of visits, the access popularity of the first document approaches 1.
[0080] In the example shown in Formula 3 above, the normalized number of visits to the first document is directly used as the access popularity of the first document.
[0081] In other examples, after normalizing the number of visits to the first document, the access popularity of the first document can also be determined using Formula 4 as follows. In this example, the access popularity of the first document can be negatively correlated with the number of visits to the first document.
[0082] behavior factor = 1.5-normalized behavior Formula 4
[0083] In Formula 4, behavior factor is the access heat of the first document, and the value range is greater than or equal to 0.5 and less than or equal to 1.5. normalized behavior is the normalized access amount of the first document, which can be referred to the introduction in Formula 3.
[0084] As can be seen from Formula 4, the higher the access amount of the first document is, the lower the access heat of the first document is. When the access amount of the first document is the lowest, the normalized access amount of the first document is 0, and the access heat of the first document is 1.5. When the access amount of the first document is the highest, the normalized access amount of the first document is 1, and the access heat of the first document is 0.5.
[0085] The above introduces the determination of the access heat of the first document. The following introduces the implementation of determining the time decay amount based on the age of the first document and the access heat of the first document.
[0086] In Formula 5, Time
[0087] In some implementations, in combination with Formula 4, the access heat of the first document and the time decay amount can be positively correlated, the greater the access heat of the first document is, the greater the time decay amount is, that is, the decay is accelerated. Conversely, the smaller the access heat of the first document is, the smaller the time decay amount is, that is, the decay is slowed down. In other words, the higher the access amount of the first document is, the smaller the time decay amount is. The higher the access amount of the first document is, the smaller the time decay amount is.
[0088] As a specific example, the time decay amount can be determined based on the age of the first document and the access heat of the first document by using Formula 5.
[0089]
[0090] In Formula 5, Time decay is the time decay amount, λ1 is a short-term decay rate. The sensitivity of emphasizing recent changes can be adjusted according to actual business requirements, such as taking a relatively large value of 0.02. Δt is the age of the first document, behavior factor is the access heat of the first document shown in Formula 4.
[0091] The embodiment takes the access data of the document as one of the factors affecting the validity of the document, quantifies the validity of the document by the user access behavior, and thus makes the validity of the document more accurate.
[0092] Optionally, after the time decay amount is obtained, the manner of determining the validity of the first document based on the sentiment score and the time decay amount can include: performing weighted summation on the sentiment score and the time decay amount to obtain a timeliness index of the first document, so as to determine the validity of the first document according to the timeliness index of the first document.
[0093] As a possible example, the higher the timeliness index, the higher the value and the stronger the validity of the first document. Conversely, the lower the timeliness index, the lower the value and the weaker the validity of the first document.
[0094] As another possible example, the timeliness index of the first document can be compared with a preset index threshold. If the timeliness index of the first document is greater than or equal to the preset index threshold, it is determined that the first document is valid; if the timeliness index of the first document is less than the preset index threshold, it is determined that the first document is invalid.
[0095] As a specific example, the timeliness index can be determined based on the sentiment score and the time decay amount by using Formula 6.
[0096] Timeliness index = a · sentiment index + β · Time decay Formula 6
[0097] wherein Timeliness index is the timeliness index, sentiment index is the sentiment score, Time decay is the time decay amount, a is the weight corresponding to the sentiment score, and β is the weight corresponding to the time decay amount.
[0098] Optionally, a and β, i.e., the weight corresponding to the sentiment score and the weight corresponding to the time decay amount, as shown in Formula 5, can be determined based on the type of the first document.
[0099] For example, when the type of the first document is a type with high time sensitivity, it means that the document content corresponding to the first document at different periods can be quite different, and at this time, the weight corresponding to the time decay amount can be increased and the weight corresponding to the sentiment score can be correspondingly reduced. When the type of the first document is a type with weak time sensitivity, it means that the document content of the first document will not change greatly in a long time, and at this time, the weight corresponding to the sentiment score can be increased and the weight corresponding to the time decay amount can be correspondingly reduced.
[0100] For example, since the application interface is closely related to technical iteration, when the first document is a document related to the application interface, the weight corresponding to the time decay amount can be adjusted to 0.7, and the weight corresponding to the sentiment score can be adjusted to 0.3.
[0101] For another example, since the theoretical document generally does not change frequently in a short time, when the first document is a theoretical document, the weight corresponding to the sentiment score can be adjusted to 0.7, and the weight corresponding to the time decay amount can be adjusted to 0.3.
[0102] The above describes the implementation of determining the effectiveness of the first document based on the sentiment score and the time decay amount. In other embodiments, the effectiveness of the first document can also be determined based on the sentiment score, the time decay amount, and the age of the first document.
[0103] Optionally, in this embodiment, the decay factor can be determined according to the age of the first document first, and then the effectiveness of the first document can be determined based on the sentiment score, the time decay amount, and the decay factor.
[0104] For example, the decay factor can be the result of inputting the age of the first document into a pre-constructed logarithmic function.
[0105] For another example, the decay factor can be the result of inputting the age of the first document into a pre-constructed linear function.
[0106] For another example, the decay factor can be the result of inputting the age of the first document into a pre-constructed exponential function.
[0107] For example, taking the exponential function as an example, the decay factor can satisfy the following formula 7.
[0108]
[0109] In formula 7, decay factor is the decay factor, Δt is the age of the first document, and λ2 is a long-term decay rate, reflecting the natural elimination period of the document. Here, the decay factor reflects the decay that is irrelevant to the content of the document and is only related to time. The longer the publication time of the document, the smaller the decay factor.
[0110] Optionally, in actual application, the value of the long-term decay rate λ2 in formula 7 can be a little smaller than the value of the short-term decay rate λ1 in formula 5.
[0111] Optionally, the value of λ2 can be different for different document types. For example, when the first document is a document related to the application interface, λ2 can be set to 0.005. For another example, when the first document is a theoretical document, λ2 can be set to 0.001.
[0112] Further, in some embodiments, after obtaining the attenuation factor, the validity of the first document can be determined according to the sentiment score, the weight corresponding to the sentiment score, the time attenuation amount, the weight corresponding to the time attenuation amount, and the attenuation factor.
[0113] For example, the timeliness index is obtained according to the sentiment score, the weight corresponding to the sentiment score, the time attenuation amount, and the weight corresponding to the time attenuation amount. Then, the timeliness index is compared with a preset index threshold, and the attenuation factor is compared with a preset attenuation factor threshold. When the timeliness index is less than the preset index threshold or the attenuation factor is less than the preset attenuation factor threshold, the first document is determined to be invalid. When the timeliness index is greater than or equal to the preset index threshold and the attenuation factor is greater than or equal to the preset attenuation factor threshold, the first document is determined to be valid.
[0114] In this way, by combining the timeliness index and the attenuation factor to determine the validity of the first document, not only can the problem that the obtained timeliness index focuses more on the content of the document be solved, but also the time can be taken as a factor affecting the validity of the document, so that the determined validity of the document is more accurate.
[0115] Optionally, if the attenuation factor is less than the preset attenuation factor threshold, it is considered that the first document has been punished in the time dimension. At this time, to avoid repeated punishment, the weight corresponding to the time attenuation amount can be reduced.
[0116] In other embodiments, the target information can also be directly input into a pre-trained model, and the validity of the first document is output by the model.
[0117] Optionally, the pre-trained model can be various machine learning models or deep learning models.
[0118] In some embodiments, a knowledge graph can also be generated. Optionally, the generated knowledge graph can also be output.
[0119] The knowledge graph includes the association relationship between the documents and the attribute information of each document. For example, the attribute information of each document includes at least one of a timeliness index, an attenuation factor, and a document type. The timeliness index is determined based on the sentiment score of the document and the time attenuation amount of the document, and the attenuation factor is determined based on the age of the document. The determination method of the timeliness index and the determination method of the attenuation factor have been described in the foregoing embodiments, and will not be described here.
[0120] Exemplarily, the documents in the knowledge base can be nodes, at least one of the timeliness index, the decay factor, the document type and the target information corresponding to the document can be a node attribute, and the association relationship between the documents can be an edge between the nodes. Exemplarily, the association relationship between the documents can be identified based on a graph neural network (GNN) or a PageRank algorithm.
[0121] The knowledge graph constructed in this embodiment covers all the documents in the knowledge base, the attribute information of each document and the association relationship between the documents. Compared with separately storing the documents and the document attributes, the knowledge graph construction manner is more comprehensive and intuitive, facilitating the staff to understand the whole process of document generation, iteration and elimination, facilitating the staff to more intuitively track and manage the documents, and improving the management efficiency.
[0122] In some examples, the association relationship between the documents can be a time sequence relationship. The time sequence association refers to that the versions of two documents are adjacent.
[0123] For example, there is a time sequence relationship between the node corresponding to the 1.0 version of the document A and the node corresponding to the 2.0 version of the document A, and thus an edge is established between the node corresponding to the 1.0 version of the document A and the node corresponding to the 2.0 version of the document A. However, if the document A further has a 3.0 version, there is no time sequence relationship between the node corresponding to the 1.0 version of the document A and the node corresponding to the 3.0 version of the document A, that is, there is no edge connecting the two nodes. However, there is a time sequence relationship between the node corresponding to the 2.0 version of the document A and the node corresponding to the 3.0 version of the document A, that is, there is an edge connecting the two nodes.
[0124] In other examples, the association relationship between the documents can be a reference relationship. The reference relationship refers to that one document references another document.
[0125] For example, the document A is an application program interface document, and the document B is a document related to some application program interfaces involved in the document A, and thus there is a reference relationship between the document A and the document B, that is, the document B references the document A. At this time, an edge can be established between the node corresponding to the document A and the node corresponding to the document B, and the edge is used to connect the two nodes. Optionally, the edge can be a directed edge, for example, the edge can be directed from the document A to the document B.
[0126] Optionally, the knowledge graph can be a graph data structure as shown in Figure 2 .
[0127] Optionally, the knowledge graph can also be a tree data structure as shown in Figure 3 .
[0128] Optionally, the knowledge graph can be stored in a graph database (such as Neo4j).
[0129] In some embodiments, the condition for generating the knowledge graph comprises at least one of the following: detecting a new document published by a user; detecting that at least one of the timeliness index and the decay factor of the document does not satisfy a preset condition; satisfying a preset period for updating the knowledge graph.
[0130] Optionally, when the new document published by the user is detected, the document version number, the document type, and the target information of the new document can be obtained, and the knowledge graph is updated according to the document version number, the document type, and the target information.
[0131] For example, the knowledge graph can be traversed to determine whether there is a node in the knowledge graph that has a time sequence relationship with the new document according to the document version number of the new document. It is determined whether there is a node in the knowledge graph that has a reference relationship with the new document according to the document type of the new document. The knowledge graph is updated according to the determination result.
[0132] In some embodiments, when at least one of the timeliness index and the decay factor of the document does not satisfy the preset condition is detected, the node in the knowledge graph that does not satisfy the preset condition can be marked, so that the knowledge base administrator or the publisher of the document corresponding to the marked node updates the document corresponding to the marked node in the knowledge base according to the marked node in the knowledge graph. After completing the update operation on the document in the knowledge base, the knowledge graph is updated based on the update operation.
[0133] The document corresponding to the marked node is the document to be updated. For example, the document to be updated includes but is not limited to a document to be deleted and a document to be iterated. The preset condition corresponding to the timeliness index is the preset index threshold in the foregoing embodiments, and the preset condition corresponding to the decay factor is the preset decay factor threshold in the foregoing embodiments.
[0134] For example, after the knowledge base administrator confirms that the document corresponding to the marked node is outdated after viewing the document, the document corresponding to the marked node is deleted from the knowledge base. At this time, in response to the deletion operation on the document corresponding to the marked node, the marked node and the edge connected to the marked node in the knowledge graph are deleted.
[0135] By marking the node in the knowledge graph that does not satisfy the preset condition, the knowledge base administrator or the publisher of the document corresponding to the marked node can re-audit the invalid document corresponding to the marked node in the knowledge base when observing that there is a marked node in the knowledge graph, thereby improving the management efficiency of the document in the knowledge base.
[0136] Optionally, the marking manner includes but is not limited to modifying the fill color of the marked node in the knowledge graph, adding a prompt mark on the edge of the marked node, etc. In this way, by marking the node corresponding to the document that does not meet the preset condition in the knowledge graph, the administrator can understand the status of the documents in the knowledge base, and can take corresponding measures in time, such as updating the content of the document corresponding to the marked node and deleting the document corresponding to the marked node.
[0137] Optionally, when the marked document is a document to be deleted, the nodes having an association relationship with the node corresponding to the document to be deleted also need to be marked in the knowledge graph to realize synchronous updating.
[0138] In this way, by marking the nodes having an association relationship with the document to be deleted in the knowledge graph, the documents having an association relationship with the deleted document can be updated in time, avoiding the risk of breaking the knowledge chain and providing a guarantee for the integrity of the documents in the knowledge base.
[0139] The following describes the scheme provided by the embodiments of the present application in combination with a specific example.
[0140] Taking the evaluation of the effectiveness of the technical document "Docker Container Network Configuration Guide" (which can be taken as an example of the first document) as an example. The target information of the document includes a timestamp (such as the publication time: January 1, 2023, and the current time: June 11, 2025), access data (such as the number of accesses: 50, the number of collections: 5, and the number of shares: 2), and comment data (such as "The document is still based on Docker v20.10, but the current latest version is v25.0, the network configuration method has changed, and it is recommended to update!"). The maximum access volume is 1000, and the minimum access volume is 10.
[0141] The part or all of the content, the access data, and the comment data included in the document are input into the preset model, and the sentiment score S of the document output by the preset model is -0.8.
[0142] Then, the sentiment score is normalized based on the above formula 1 to obtain the normalized sentiment score sentiment index , the sentiment score sentiment index satisfies:
[0143] The age Δt of the document is determined according to the publication time and the current time of the document, and the age Δt of the document satisfies: Δt = 365 x 2.5 ≈ 912 days.
[0144] The access volume current activity of the document is determined based on the access data of the document by using the above formula 2. The access volume currentactivity satisfies: current activity = x·visits + y·favorites + z·shares = 1.0 x 50 + 1.5 x 5 + 2.0 x 2 = 50 + 7.5 + 4 = 61.5.
[0145] According to the access amount, the maximum access amount and the minimum access amount of the document, the access heat behavior of the document is determined by using the above formula 2, formula 3 and formula 4. factor The access heat behavior of the document is factor satisfies:
[0146]
[0147] Based on the age and the access heat of the first document, the time decay amount Time is determined by using the above formula 5. decay The time decay amount Time is decay satisfies:
[0148] Based on the sentiment score and the time decay amount, the timeliness index Timeliness is determined by using the above formula 6. index The timeliness index Timeliness is index satisfies: Timeliness index = a·sentiment index + b·Time decay = 0.6 x 0.1 + 0.4 x 0 = 0.06.
[0149] According to the age of the first document, the decay factor decay is determined by using the above formula 7. factor The decay factor decay is factor satisfies:
[0150] Taking the preset index threshold as 0.3 and the preset decay factor threshold as 0.3 as an example, the timeliness index is less than the preset index threshold, which indicates that the user feedbacks negative to the document and the access heat is extremely low, and it is determined that the short-term of the document is almost zero. The decay factor is less than the preset decay factor threshold, which indicates that the document has been seriously naturally aged due to technical iteration. Further, it is determined that the document is invalid.
[0151] Optionally, the node corresponding to the document in the knowledge graph is further marked, which can be marked as "technology elimination document, need to be rewritten completely". At the same time, the author can also be reminded that "the user points out that the document version is out of date, and the content needs to be updated".
[0152] Further, the information of the node corresponding to the document in the knowledge graph is updated, such as attribute information, association relationship and the like.
[0153] There is also provided in the embodiments a document validity determining apparatus for implementing the above embodiments and preferred embodiments, which has been described and will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and contemplated.
[0154] The embodiments provide a document validity determining apparatus, as shown in Figure 4 The apparatus comprises:
[0155] The apparatus comprises:
[0156] The apparatus comprises:
[0157] In some optional embodiments, the determining module 402 comprises:
[0158] The apparatus comprises:
[0159] In some optional embodiments, the timestamp of the first document comprises a publishing time; and the second determining sub-module comprises:
[0160] The apparatus comprises:
[0161] In some optional embodiments, the third determining unit comprises:
[0162] The apparatus comprises:
[0163] In some optional embodiments, the third sub-unit is specifically configured to:
[0164] The validity of the first document is determined based on the sentiment score, the weight corresponding to the sentiment score, the time decay amount, the weight corresponding to the time decay amount, and the age of the first document.
[0165] In some optional embodiments, the weight corresponding to the sentiment score and the weight corresponding to the time decay amount are determined based on the type of the first document.
[0166] In some optional embodiments, the apparatus further comprises:
[0167] The generation module is configured to generate a knowledge graph, wherein the knowledge graph comprises an association relationship between documents and attribute information of each document, and the attribute information of each document comprises at least one of a timeliness index, a decay factor, and a document type; wherein the timeliness index is determined based on a sentiment score of a document and a time decay amount of the document, and the decay factor is determined based on the age of the document.
[0168] In some optional embodiments, the condition for generating the knowledge graph comprises at least one of the following:
[0169] detecting a new document published by a user; detecting that at least one of the timeliness index and the decay factor of the document does not meet a preset condition; and meeting a preset period for updating the knowledge graph.
[0170] Further function descriptions of the above-mentioned modules and units are the same as those of the corresponding embodiments, and will not be repeated here.
[0171] The document validity determination apparatus in the present embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory executing one or more software or fixed programs, and / or other devices that can provide the above functions.
[0172] The present embodiment further provides a computer device with the above-mentioned Figure 4 document validity determination apparatus.
[0173] Please refer to Figure 5 , Figure 5 is a structural schematic diagram of a computer device provided by an optional embodiment of the present application, as Figure 5As shown, the computer device includes one or more processors 501, memory 502, and interfaces 505 for the various components. The various components communicate over one or more busses 504 that are associated with the different interfaces 505. Various components can be mounted on a common motherboard or in other manners as appropriate. The processor 501 can process instructions for execution within the computer device, including instructions stored in the memory 502 or elsewhere within the computer device. The graphics information of the instructions can be processed to display graphical information for a GUI on an external input / output device, such as a display device coupled to the interface 505. In some alternative implementations, multiple processors and / or multiple buses can be used, as appropriate, along with multiple memories and types of memory. Also, multiple computer devices can be connected, with each computer device providing portions of the necessary operations (e.g., as a server array or a group of blade servers, or multiple processors). Figure 5 The processor 501 is taken as an example in the embodiment.
[0174] The processor 501 can be a central processing unit, a network processor, or a combination thereof. The processor 501 can further include a hardware chip. The hardware chip can be an application specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device can be a complex programmable logic device, a field programmable logic device, a general array logic, or any combination thereof.
[0175] The memory 502 stores instructions that are executable by the at least one processor 501, so as to enable the at least one processor 501 to perform the method shown in the above embodiment.
[0176] The memory 502 can include a program storage area and a data storage area. The program storage area can store an operating system, application programs required by at least one function, and the like. The data storage area can store data created according to the use of the computer device, and the like. In addition, the memory 502 can include a high-speed random access memory, and can further include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some alternative implementations, the memory 502 can optionally include a memory that is remotely arranged with respect to the processor 501, and these remote memories can be connected to the computer device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0177] The memory 502 can include a volatile memory, such as a random access memory, and can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid state disk. The memory 502 can further include a combination of the above-mentioned kinds of memories.
[0178] The computer device further includes a communication interface 503 for communication of the computer device with other devices or communication networks.
[0179] The embodiments of the present application further provide a computer readable storage medium, and the method according to the embodiments of the present application can be implemented in hardware, firmware, or recorded in a storage medium, or stored in a remote storage medium or a non-transitory machine readable storage medium and downloaded to a local storage medium through network and stored in the local storage medium, so that the method described herein can be processed by such software on a storage medium using a general purpose computer, a special purpose processor, or programmable or special hardware. The storage medium can be a disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid state disk, etc. Further, the storage medium can also include a combination of the above-mentioned memories. It can be understood that the computer, the processor, the microprocessor controller, or the programmable hardware includes a storage component that can store or receive software or computer code, which, when accessed and executed by the computer, the processor, or the hardware, implements the method shown in the above embodiments.
[0180] In addition to the above computer device and computer readable storage medium, the embodiments of the present application can also be a computer program product, which includes computer program instructions, which, when executed by a processor, causes the processor to perform the steps of the sound source positioning method provided by any embodiment of the present application.
[0181] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of the present application, including an object-oriented programming language, such as Java, C++, etc., and a conventional procedural programming language, such as "C" language or similar programming language. The program code can be executed entirely on a user computing device, partially on a user device, as an independent software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0182] It should be noted that in this paper, the relationship terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitation, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.
[0183] The foregoing detailed description of the technology has been presented for purposes of illustration and description. Various modifications to the description
[0184] embodiments will be readily apparent to those of ordinary skill in the art, and it is intended to be covered by the following claims, as not to depart from the spirit and scope of the present disclosure.
[0185] It will be apparent to those skilled in the art that various modifications and variations can be made in the present disclosure without departing from the spirit or scope of the disclosure. Thus, it is intended that the present disclosure cover the modifications and variations of this disclosure provided they come within the scope of the appended claims and their equivalents.
[0186] It is intended, therefore, that the present disclosure be considered in all its
[0187] embodiments disclosed herein, and it is intended that the present disclosure be considered in the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method of determining the validity of a document, characterized by, The method comprises: obtaining target information, the target information comprising at least two of part or all contents included in a first document, a timestamp of the first document, access data of the first document by a user, and comment data of the first document by the user; determining validity of the first document based on the target information.
2. The method of claim 1, wherein, The determining of the validity of the first document based on the target information comprises: inputting the part or all contents included in the first document, the access data of the first document by the user, and the comment data of the first document by the user into a preset model, and outputting a sentiment score of the first document by the preset model; determining the validity of the first document based on the sentiment score, the timestamp of the first document, and the access data of the first document by the user.
3. The method of claim 2, wherein, The timestamp of the first document comprises a publishing time. The determining of the validity of the first document based on the sentiment score, the timestamp of the first document, and the access data of the first document by the user comprises: determining a difference between a current time and the publishing time as an age of the first document; determining a time decay amount based on the age and an access heat of the first document, the access heat being determined based on the access data of the first document and access data of a second document, the second document being a document other than the first document stored in a storage space; determining the validity of the first document based on the sentiment score and the time decay amount.
4. The method of claim 3, wherein, The determining of the validity of the first document based on the sentiment score and the time decay amount comprises: determining the validity of the first document based on the sentiment score, the time decay amount, and the age of the first document.
5. The method of claim 4, wherein, The determining of the validity of the first document based on the sentiment score, the time decay amount, and the age of the first document comprises: determining the validity of the first document based on the sentiment score, a weight corresponding to the sentiment score, the time decay amount, a weight corresponding to the time decay amount, and the age of the first document.
6. The method of claim 5, wherein, The weight corresponding to the sentiment score and the weight corresponding to the time decay amount are determined based on a type of the first document.
7. The method according to any one of claims 3-6, characterized in that, The method further comprises: generating a knowledge graph, wherein the knowledge graph comprises an association relationship between documents and attribute information of each document, the attribute information of each document comprising at least one of a timeliness index, a decay factor, and a document type; wherein the timeliness index is determined based on a sentiment score of the document and a time decay amount of the document, and the decay factor is determined based on an age of the document.
8. The method of claim 7, wherein, The condition for generating the knowledge graph comprises at least one of the following: detecting a new document published by a user; detecting that at least one of a timeliness index and a decay factor of a document does not meet a preset condition; satisfying a preset period for updating the knowledge graph.
9. An apparatus for determining validity of a document, characterized by The apparatus comprises: An acquisition module is configured to acquire target information, the target information including at least two of the following: part or all of the content included in a first document, a timestamp of the first document, access data of the first document by a user, and comment data of the first document by the user. A determination module is configured to determine validity of the first document based on the target information.
10. A computer device, comprising: The method comprises the following steps: A memory and a processor are in communication connection with each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the method in any one of claims 1 to 8.