The invention discloses a method, a
system, a medium and equipment for extracting unified aggregated text content of all types of files, and relates to the technical field of file
processing, the method comprises the following steps:
system initialization and configuration, signature
verification, extraction judgment, proxy forwarding,
content extraction, and result storage and return. A single and stable service
entry point (Endpoint) and a standardized API (Application Program Interface) specification are provided for all upper-layer businesses; through a standard S3 signature
verification mechanism, the security and reliability of each call are ensured; in combination with an AK / SK
system, unified
authority control,
flow limitation and post auditing can be conveniently carried out; a service-oriented architecture is adopted, and the extraction capability of different types of files is provided by an independent micro-
service module. When a new file type needs to be supported, only a new
content extraction module needs to be developed and registered into the mapping table, a main process and an upper layer service do not need to be changed, and the system expansibility is extremely high. And in combination with
file size judgment, executing logic of combining synchronous extraction and asynchronous extraction of the file.