Connecting SharePoint to Knowledge Management
The SharePoint knowledge integration syncs content from your SharePoint Online sites into the Rezolve.ai Knowledge Base. The Virtual Agent can then use that content to answer questions and cite the source document. Agents can also use it as a reference.
The integration connects through the Microsoft Graph API using an Azure (Entra ID) app registration that you own. You can sync:
| Content | What gets ingested |
|---|---|
| Drives (document libraries) | Every supported file in the library, including all sub-folders |
| Folders | Every supported file in the folder and its sub-folders |
| Site Pages | The text of modern SharePoint pages (SitePages/*.aspx), either all pages of a site or selected pages |
| Sub-sites | Content of the selected sub-sites, either the whole sub-site or specific libraries, folders, pages, and lists within it |
| Lists | Each list item becomes a knowledge entry. You choose which columns are searchable. |
This page covers the SharePoint-specific setup. For the wizard, sync status, bounced files, and other functions that all integrations share, see Knowledge Integrations.
Prerequisites
- SharePoint Online (Microsoft 365 commercial cloud). SharePoint Server (on-premises) and the national clouds (GCC High, DoD, China/21Vianet) are not supported.
- An Azure / Entra ID administrator who can register an application and grant admin consent.
- For
Sites.Selectedonly: a SharePoint or Global administrator to grant the app access to each site. - Knowledge Management admin access in the Rezolve.ai console.
Step 1: Register an Azure application
- In the Azure portal, open App registrations and click + New registration.
- Enter a name, for example
Rezolve.ai SharePoint Ingestion. Keep the default single-tenant account type, then click Register. - On the Overview page, copy these two values:
- Directory (tenant) ID, which is the Azure Tenant ID in Rezolve.ai
- Application (client) ID, which is the Client ID in Rezolve.ai
- Open Certificates & secrets → Client secrets → + New client secret. Enter a description, choose an expiry, and click Add.
- Copy the secret Value right away. Azure shows it only once.
Only client secret authentication is supported. Certificate-based authentication is not. Note the secret's expiry date: when it expires, syncs fail until you update the connection with a new secret.
For a screenshot walkthrough, see How to Register an Azure Application for SharePoint Ingestion.
Step 2: Grant Microsoft Graph permissions
Rezolve.ai only reads from SharePoint. Choose one of the two permission models below. Both are Application permissions (not Delegated) and both need admin consent.
| Permission model | Access granted | When to use |
|---|---|---|
| Sites.Read.All | Read access to every site in the tenant | Simplest setup, or when you'll sync many sites |
| Sites.Selected (recommended) | Read access only to sites you explicitly grant | Least-privilege setup. Access is limited to the sites you sync. |
The Rezolve.ai app doesn't need the Files.Read.All, Sites.ReadWrite.All, or Sites.FullControl.All Graph permissions. Don't grant them to the ingestion app. The one exception is a site-level fullcontrol grant under Sites.Selected, which is only needed for permission inheritance.
Option A: Sites.Read.All
- In your app registration, open API permissions → + Add a permission → Microsoft Graph → Application permissions.
- Search for and select Sites.Read.All, then click Add permissions.
- Click Grant admin consent for <your directory>.
That's all. The app can now read every site.
Option B: Sites.Selected (per-site access)
With Sites.Selected, consent alone gives the app no access. You must also grant it read access on each site you want to sync.
-
In your app registration, add the Microsoft Graph → Application → Sites.Selected permission and click Grant admin consent.
-
Find the ID of each site. For example, in Graph Explorer:
GET https://graph.microsoft.com/v1.0/sites/yourcompany.sharepoint.com:/sites/HRThe
idin the response looks likeyourcompany.sharepoint.com,<guid>,<guid>. -
Grant the Rezolve.ai app read access to the site. Run this as an administrator with
Sites.FullControl.All, using Graph Explorer, PowerShell, or a separate admin app. Do not use the Rezolve.ai ingestion app for this step.curl --location 'https://graph.microsoft.com/v1.0/sites/<SITE_ID>/permissions' \
--header 'Authorization: Bearer <ADMIN_ACCESS_TOKEN>' \
--header 'Content-Type: application/json' \
--data '{
"roles": ["read"],
"grantedToIdentities": [
{
"application": {
"id": "<CLIENT_ID>",
"displayName": "<APP_DISPLAY_NAME>"
}
}
]
}'Placeholder Value <SITE_ID>The site ID from step 2 <ADMIN_ACCESS_TOKEN>A token for an admin identity that holds Sites.FullControl.All<CLIENT_ID>The Application (client) ID of the Rezolve.ai ingestion app <APP_DISPLAY_NAME>The display name of the Rezolve.ai ingestion app -
Repeat step 3 for every site collection you want to sync.
- Classic sub-sites under the site's URL, such as
/sites/HR/Benefits, are covered by the grant on the parent site. - Sites that are separate site collections need their own grant, even when they appear related, for example hub-associated sites such as
/sites/HR-Policies.
- Classic sub-sites under the site's URL, such as
Permission inheritance reads each file's sharing permissions. SharePoint returns the complete list only to site owners, so grant the fullcontrol role instead of read on sites where you enable Inherit Permissions from SharePoint. The write role is not enough. The app still only reads content.
Don't add Sites.Read.All or Files.Read.All to an app that uses Sites.Selected. Those permissions give access to every site and override the per-site restriction.
To grant access with PnP PowerShell instead of the Graph API, see Sites.Selected - Enabling API Permissions for SharePoint Sites Crawl with PowerShell Script.
If the app can't read a site, library, or page, Rezolve.ai skips it without reporting an error. When a sync finishes but content is missing, check the permission grant first. See Troubleshooting.
Step 3: Create the SharePoint integration
- In the Rezolve.ai console, go to Knowledge Management → Knowledge Integrations.
- Click Add Integrations and select SharePoint.
3a. Connection setup
Enter a Title for the integration. Then either choose an existing SharePoint connection (Choose From Existing Connections) or select Create New Connection and enter these details:
| Field | Description | Example |
|---|---|---|
| Name | A descriptive name for the connection | Contoso SharePoint - HR |
| Site URL | Full URL of the SharePoint site to sync | https://yourcompany.sharepoint.com/sites/HR |
| Azure Id | Directory (tenant) ID from Step 1 | 72f988bf-86f1-41af-91ab-2d7cd011db47 |
| Client Id | Application (client) ID from Step 1 | bf7f3a64-7c3c-4c71-91a2-2b50fd0a7f2c |
| Client Secret | Client secret value from Step 1 (not the secret ID) | ~2TjQ8... |
Site URL rules:
- Use the full URL of a site under
/sites/or/teams/. Any other URL form is treated as the tenant root site. - Use the
*.sharepoint.comaddress. Custom or vanity domains are not resolved. - The Site URL can't be edited after the integration is created. To sync a different site, create a new integration.
When editing a connection later, you can replace the Client Id and Client Secret, for example when the secret expires.
3b. Content selection
Content is organized into one or more content sets. Each content set has its own selection and access settings. In each content set, browse the site tree and tick what to sync. You don't need to type any IDs.
| Item in the tree | What it syncs |
|---|---|
| Drives (document libraries) | The whole library, crawled recursively. Expand a drive to pick individual folders instead. The Preservation Hold Library is never synced. |
| Folders | The folder and all of its sub-folders. A folder inside a ticked drive is already covered and isn't processed twice. |
| Pages | Tick the Pages folder to sync all site pages, including pages published later. Expand it to pick individual pages; pages load 50 at a time, and Load More shows the rest. Individually picked pages stay a fixed list. Unticking any page switches the content set to an individual page list. |
| Sub-sites | Tick a sub-site to include its content, or expand it to pick its drives, folders, pages, lists, and nested sub-sites individually. |
| Lists | Shown when the SharePoint List switch in the content set's ⋮ menu is on. See SharePoint lists. |
Restricting access
Open the content set's ⋮ menu and turn on Restrict Access. Then choose either or both of:
- Restrict by Audience: pick one or more Audience(s) that can get answers from this content set.
- Inherit permissions from SharePoint: restrict each item to the SharePoint users who can access it. See Permission inheritance.
Without Restrict Access, content is available to everyone who uses the Virtual Agent.
SharePoint lists
To sync SharePoint lists, open the content set's ⋮ menu and turn on SharePoint List. This switch appears for users with the Knowledge admin role. The site's lists then appear in the tree. Only visible custom lists that have visible columns are offered.
Tick a list and expand it to choose its fields:
| Setting | Purpose |
|---|---|
| Contextual Fields | The columns whose values become the searchable text of each item. Select at least one. |
| Meta Fields | Columns stored as metadata alongside each item. Remove any you don't need. |
Each list row becomes one knowledge entry that links to the item's display form (DispForm.aspx?ID=...).
3c. Configuration
Under Sync Details, set the Frequency (daily, weekly, or monthly), Start Time, Timezone, and Schedule Start Date. Optionally turn on Subscribe to Sync Reports and add User Subscribers. Then click Save. See Knowledge Integrations → Step 3: Configuration.
What gets ingested
Supported file types
| Category | Extensions |
|---|---|
| Documents | .pdf, .doc, .docx, .txt, .md |
| Presentations | .ppt, .pptx |
.msg | |
| Archives | .zip, .rar, .7z (supported files inside are extracted) |
| Pages | Modern site pages (.aspx) |
Excel (.xls, .xlsx), CSV, images, video, and audio files in document libraries are not ingested. They appear as bounced files with an unsupported file type reason.
Limits
| Limit | Default |
|---|---|
| Maximum file size | 25 MB per file |
| Archive contents | Up to 10 files and 50 MB uncompressed per archive |
| Retries per file | A file that fails 3 times, or fails for a reason that won't change, is skipped on later syncs until it's resolved in Bounced Files |
Site page content
- Only the text web parts of a page are ingested. Images, embedded videos, lists, and other non-text web parts are ignored.
- Links in page text are kept and made absolute.
- A page with no text is skipped. It isn't treated as a failure.
Metadata
Each ingested document keeps:
- its name
- its SharePoint URL, which the Virtual Agent shows as the source link
- its last-modified time
- its SharePoint column values, including columns inherited from parent folders
Draft or checked-out versions aren't filtered out. Whatever the app can read in the library is ingested. Keep drafts in a location you don't sync.
Sync behavior
| Behavior | Details |
|---|---|
| First sync | Full ingestion of everything selected |
| Scheduled / Sync Now | Incremental. Each run compares the selected locations with what's already ingested, then adds new items, re-ingests items modified since the last sync, and removes items that no longer exist in SharePoint. |
| Moves and renames | Treated as a delete at the old location plus an add at the new one |
| Removing a source from a content set | When you save, the content from the deselected source is removed from the Knowledge Base. If you deselect a drive or sub-site, a sync also starts automatically. |
| Other content changes | Newly added sources are synced on the next scheduled sync, or right away with Sync Now |
| Stop Sync | Stops future syncs. Content already ingested stays in the Knowledge Base. |
| Delete integration | Removes the configuration and all content it ingested |
| Throttling | When Microsoft Graph throttles requests (HTTP 429) or has temporary errors, requests are retried automatically, respecting the Retry-After header |
Each run enumerates every selected location, so very large libraries take longer to sync even when little has changed. Schedule large syncs outside business hours.
Permission inheritance
When Inherit Permissions from SharePoint is enabled, the permissions below are recorded with each item. The Virtual Agent then only returns content to users who have access to it.
| Content | Permissions recorded |
|---|---|
| Files in drives and folders | The SharePoint groups and Entra ID groups that have read, edit, or owner access to the file |
| Site pages | Site-level access only. Per-page unique permissions are not captured. |
| List items | The list's configured audience column, if one is set |
Access granted directly to individual users, rather than through a group, isn't captured. For accurate access control, manage SharePoint permissions through groups.
Troubleshooting
| Symptom | Likely cause | Resolution |
|---|---|---|
| Sync succeeds but a site, library, or page is missing | The app has no access to it. This is common with Sites.Selected when a separate site collection (for example a hub-associated site) wasn't granted. | Grant read access to that site (Step 2, Option B), then click Sync Now |
| All syncs fail after working earlier | The client secret has expired | Create a new secret in Azure and update the connection's Client Secret |
| Authentication error when saving the connection | Wrong tenant ID, client ID, or secret. The secret ID was entered instead of the secret value. | Re-copy the values from the app registration |
| Content from the tenant root site appears instead of the expected site | The site URL isn't in /sites/... or /teams/... form | Use the full https://<tenant>.sharepoint.com/sites/<name> URL |
| Site URL not found | A custom or vanity domain was used | Use the *.sharepoint.com URL |
| File listed in Bounced Files as size limit exceeded | File is over 25 MB | Split or compress the file, or exclude it |
| File listed as unsupported file type | For example, Excel, CSV, or image files | Convert to a supported format (for example, PDF) if the content is needed |
| Permission inheritance is on but users can't get answers they should see, or the audience looks incomplete | With Sites.Selected, the app has only the read role, so file permissions come back incomplete | Change the site grant to fullcontrol, then click Sync Now |
| Scanned PDF ingested with little or no text | Text in images isn't read | Use a PDF with a text layer, for example by running OCR in Acrobat before uploading |
| Page ingested with missing content | The content is in image, embed, or other non-text web parts | Put the key information in text web parts |
| A file stopped retrying | It failed 3 times | Fix the file in SharePoint, resolve it in Bounced Files, then click Sync Now |
Download the CSV log for any sync run from the Sync Status tab. See Downloading Sync Logs.
Resources
| Resource | Link |
|---|---|
| Register an application in Entra ID | Microsoft Learn |
| Graph permissions reference (Sites) | Microsoft Learn |
Selected permissions in SharePoint (Sites.Selected) | Microsoft Learn |
| Create site permission (Graph API) | Microsoft Learn |
| Microsoft Graph SharePoint API | Microsoft Learn |
Next Steps
- Manage syncs, logs, and bounced files in Knowledge Integrations
- Set up audience targeting for SharePoint content
- Monitor usage with analytics