aws-textract
Lets the project run AWS Textract over a stored image or PDF to extract its text. Optionally, the same call passes the text to AWS Comprehend Medical.
| Type | Project feature |
Value in features | aws-textract |
| Change with | PATCH /v1/projects/me/features (replaces the whole list) |
| Who can change it | Project admin. Any practitioner can read the list. |
| On for new projects | No |
| When off | POST /fhir/R4/:type/:id/$aws-textract answers 400 |
| Related feature | aws-comprehend |
What it enables
POST /fhir/R4/:type/:id/$aws-textract, where :type is Media or DocumentReference and the resource points at a stored file (a Binary).
curl --request POST \
--url "https://api.sandbox.ovok.com/fhir/R4/Media/${MEDIA_ID}/\$aws-textract" \
--header "Authorization: Bearer ${OVOK_TOKEN}" \
--header 'Content-Type: application/json' \
--data '{}'
| Body | Result |
|---|---|
{} or no body | Textract only. |
{"comprehend": true} | Textract, then Comprehend Medical entity detection on the extracted text. |
The call waits for the analysis, which can take up to about two minutes, then returns the raw Textract output and writes the output into your project as a new Binary and Media resource.
Turn it on
curl --request PATCH \
--url 'https://api.sandbox.ovok.com/v1/projects/me/features' \
--header "Authorization: Bearer ${OVOK_TOKEN}" \
--header 'Content-Type: application/json' \
--data '{"features":["bots","cron","email","transaction-bundles","websocket-subscriptions","aws-textract"]}'
Start from the list GET /v1/projects/me/features returns, and add aws-textract to it.
What callers see when it is off or not possible
| Condition | Status | Message |
|---|---|---|
| 1. Feature is off | 400 | AWS Textract not enabled |
| 2. The platform does not store files in S3 | 400 | AWS Textract requires S3 storage |
3. :type is not Media or DocumentReference | 400 | A message naming the unsupported type |
Gotchas
- Your documents go to AWS. The file is analysed by AWS services, and the extracted text is stored as new resources in your project. Treat the output as the same sensitivity as the source.
- It is a paid AWS service. Textract and Comprehend Medical are billed by usage. Do not call the operation in a loop or from public code.
- The call is synchronous. A large document can run into the request time limit of your client or network. Use a generous timeout and plan for a retry.
aws-comprehendis not what turns on Comprehend here. Thecomprehendoption is governed by this feature. See aws-comprehend.- It can be unavailable even when the feature is on. The platform must store files in S3. In the sandbox the operation currently answers
400 AWS Textract requires S3 storagewith the feature on, so you cannot build on it there. The checks run in the order above, so a wrong:typeis only reported once the first two pass. - Projects do not inherit it. Turn it on in each project that needs it.